VLDB 2026 Research / reviewers in the wild / expert
Udayan Ganguly
dblp:179/2695
· DBLP profile ↗
34ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0002-1498-5993ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 12 since 2021Systems, architecture and hardware · 13 · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D Monolithic Integrated Indium Tin Oxide-Silicon Hybrid Leaky Integrate and Fire Neuron
Harshitha Gangu, Aakash Deshpande, Ranie S. Jeyakumar, Sahil Rakesh Wani, Abhishek Kadam, Udayan Ganguly, Laxmeesha Somappa, Veeresh Deshpande |
ISCAS | 6 |
| 2025 | Band to Band Tunneling-Based Low Power and Low Area Tunable Spike Delay ElementabstractBio-inspired axonal and dendritic delay-based spiking neural network algorithms perform spatiotemporal pattern recognition efficiently within feed-forward networks, making complex and suboptimal recurrent neural network structures unnecessary. Including trainable dendritic or axonal delays in feed-forward neural networks reduces neural network complexity and improves classification performance significantly. However, generating tunable low-power hardware spike delays of the biological timescale (few μs to ms ) without adding an extra penalty on the area has been challenging over the years. We present a novel band-to-band tunneling-based tunable delay element for spiking neural network hardware. The proposed low-power core delay element capable of providing spike delays of up to 0.4 ms (without explicit capacitance) consumes an area of 50 μm2in GF45RFSOI technology with a peak power of 320 nW, which is the lowest among state-of-the-art spike delay generation circuits. Moreover, the order of the spike delay can be extended to 25 ms by adding an explicit on-chip capacitance of 500 fF. Abhishek Kadam, Shreyas Deshmukh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 5 |
| 2025 | A Neuromodulation-based Spiking Neural Network using ReRAM ArrayabstractThis work proposes a neuromodulation-inspired spiking neural network using a ReRAM memory. A stashing-merging algorithm is realized to mimic the inherent neuromodulation in humans. While traditional pruning methods remove redundant parts of the network, stashing excludes well-trained neurons while training and restores all neurons at the end of training. This approach exhibits energy-efficient training in the context of a spiking neural network (SNN) since well-trained neurons can be easily identified using the spike count. The idea is validated using a ReRAM-based SNN with 10 conductance levels and performs close to a traditional artificial neural network (ANN) on an MNIST classification workload. Nirmal Shah, Jayatika Sakhuja, Udayan Ganguly, Sandip Lashkare, Laxmeesha Somappa |
ISCAS | 3 |
| 2025 | A Hardware-Software Co-Design Platform to Evaluate SNN Workloads for ReRAM-based IMCabstractResistive random access memory (ReRAM) based analog in-memory-compute (IMC) coupled with spiking neural networks (SNN) offers a promising solution to implement efficient matrix multiplication. This work presents an ARM Cortex-based ReRAM IMC for rapid SNN workload evaluation. While the software flexibility and the scheduling are provided by the ARM processing system (PS), the programmable logic (PL) provides a scalable interface to the ReRAM array through mixed-signal digital-to-analog converters (DAC). A prototype system is presented using a Zynq 7000 SoC comprising an ARM PS and PL infrastructure. Custom 8x8 ReRAM array along with row and column DACs and leaky-integrate and fire (LIF) neurons are implemented to realize the end-to-end system. A use-case of a stashing-based MNIST classification task is demonstrated using the prototype system. Nirmal Shah, Jayatika Sakhuja, Udayan Ganguly, Sandip Lashkare, Laxmeesha Somappa |
ISCAS | 3 |
| 2025 | A 0.93 nW/node Ultra-Low Power Oscillatory Neural Network using BTBT-based OscillatorsabstractCombinatorial optimization problems (COPs), when addressed using traditional von Neumann computers, demand significant computational power and substantial area as problem dimensionality increases. Hardware-based solvers, particularly those employing coupled oscillator networks to mimic Ising machines, have been explored as alternatives. However, conventional CMOS-based solutions face limitations in terms of power consumption and area. In this work, we propose a low-power 8-node oscillatory neural network using a band-to-band-tunneling-based ring oscillator in GF45RFSOI technology. This ultra-low power and low-area design efficiently solves COPs without an external perturbation signal. We use the intrinsic noise of BTBT-based oscillators to augment the phase synchronization among the coupled oscillators. The proposed system allows configurable all-to-all connectivity between ring oscillator nodes through cross-coupled capacitors. We demonstrate the system’s efficacy in solving multiple vector graph coloring problems, achieving an average power consumption of 0.93 nW (105× lower than state-of-the-art) per oscillator with a supply voltage of 1.8 V. Abhinav Thaduri, Abhishek Kadam, Laxmeesha Somappa, Udayan Ganguly, Maryam Shojaei Baghini |
ISCAS | 4 |
| 2025 | Analog and Temporary On-chip Memory for ANN Training and InferenceabstractOn-chip training at the edge becomes a primary requisite for real-time and security-sensitive artificial neural network (ANN) applications. In-memory computation (IMC) techniques have been proposed to facilitate data-intensive computational operations in ANNs. IMC-based multiply-accumulate (MAC) accelerates ANN training but suffers from significant communication overhead between the MAC engine and the off-chip storage for the intermediate data. This article proposes an analog temporary on-chip memory (ATOM) to store this intermediate data during ANN training. The ANN training architecture with the proposed ATOM has two significant advantages. First, the energy required to store intermediate data is scaled down by \(\sim\) 40 \(\times\) due to the on-chip and analog nature of the memory. Second, the proposed architecture avoids power and area-consuming analog-to-digital converters (ADCs) between neural network stages. The ATOM cell measurements are carried out from 20 fabricated chips, and the impact of ATOM characteristics on ANN system performance accuracy is analyzed. This article shows significant latency improvement of \(\sim\) 9 \(\times\) and area savings of \(\sim\) 5 \(\times\) for intermediate data storage compared to the on-chip SRAM during ANN training’s forward and backward pass operations. An improvement in the area and latency will be beneficial to instrument the area- and energy-efficient hardware system for on-chip ANN applications. Shreyas Deshmukh, Raghav Singhal, Shruti Landge, Vivek Saraswat, Anmol Biswas, Abhishek Kadam, Ajay Kumar Singh, Sreenivas Subramoney, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ACM J. Emerg. Technol. Comput. Syst. | 11 |
| 2024 | QFALT: Quantization and Fault Aware Loss for Training Enables Performance Recovery with Unreliable WeightsabstractDeep Learning is the dominant method for classification and pattern recognition tasks. Traditionally, this entails running large, high-precision models on the cloud. However, there is an increasing demand for efficient, lightweight inference using compact models on low-power devices for edge Artificial Intelligence (AI) purposes. This necessitates highly quantized models. The task of effectively quantizing a deep-learning model is not trivial and many approaches have been proposed. In this paper, we propose a regularization-based technique to perform quantization-aware training. We validate our method by training Fully connected and Convolutional Neural Networks on the Fashion-MNIST, CIFAR-10, and CIFAR-100 datasets, showing under 1% degradation in classification accuracy for 2-bit and 3-bit quantized weight models compared to the 32-bit floating-point baseline. We also show the flexibility of our regularization function to train the network in a fault-aware manner, such that the degradation in performance caused by the presence of stuckat-0 faults making certain weight states unreachable is effectively halved, even for a high level of faulty bits (10%) and low level of quantization (3-bit). This opens up possibilities of improved fault and variability-aware training for low-power and low bit-precision neuromorphic devices. Anmol Biswas, Udayan Ganguly |
IJCNN | 2 |
| 2024 | Novel SRAM based Temporary Memory for PVT Variation Tolerant Analog In-Memory ComputingabstractAnalog in-memory computing (IMC) techniques have played a significant role in vastly improving the throughput, energy efficiency, and area of on-chip ML inference engines, breaking through the Von-Neumann memory bottleneck. However, the innate susceptibility to process, voltage and temperature (PVT) variations in effect restricts analog computing to ML inference applications of low to moderate complexity. In this work, we present a novel PVT variation tolerant SRAM based temporary memory (STEM) for analog in-memory computing. The proposed technique also allows unit cells to operate at ultralow currents, thereby enabling better energy efficiency. Monte Carlo simulations show that the proposed temporary memory array achieves a unit cell current variability (σ/μ) of 2% which is 5× better than that of conventional 8T unit cell. System level simulations of a 64 × 64 macro show that the proposed STEM achieves near baseline classification accuracy on FMNIST dataset using pre-trained weights, without chip-specific training. Sivakumar Elangovan, Porus Vangala, Yeshwanth Sunnapu, Khalid Shaikh 0001, Udayan Ganguly, Maryam Shojaei Baghini |
ISCAS | 5 |
| 2024 | A Compact Low Power Multi-mode Spiking Neuron using Band to Band TunnelingabstractEfficient and compact neurons with low power consumption are crucial when designing large-scale spiking neural networks (SNNs) for hardware implementation. Many architectures in the literature showcase different spike patterns associated with biological neurons. However, using bulky capacitors to generate the different time constants related to complex neuron patterns makes these circuits area inefficient. This paper presents a band-to-band-tunneling (BTBT) based energy-efficient and compact neuron capable of producing various spike patterns. The BTBT region’s extremely low current enables different time constants while eliminating the need of bulky capacitors. The circuit is based on the Izhikevich neuron model. The proposed circuit is designed in Silicon on Insulator technology to exhibit important firing patterns observed in the biological cortex, viz. regular spiking, fast-spiking, and chattering, and it is fine-tuned for efficient operation at low subthreshold voltages. This circuit utilizes only 129 μm2area and consumes only 6.7 fJ energy per spike ( approximately 40% lower area and energy per spike than state-of-the-art multi-mode neurons) in G45RFSOI technology. Abhishek Kadam, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 5 |
| 2024 | A Compact 140nW/input Winner-Take-All Circuit for Spiking Neural NetworksabstractSolving classification problems using Spiking Neural Networks (SNNs) involves determining the most active neuron in the output layer. Scalable, low-power and low-area hardware solutions for such decision-making are vital for neuromorphic edge applications to meet power and space constraints. In this work, we propose a low-power, compact Winner-Take-All (WTA) circuit, a multi-input multi-output dynamic threshold comparator that simultaneously compares multiple analog voltage inputs and provides a one-hot-encoded digital output vector indicating the result of the classification. The design eliminates the need for cascading and a dedicated feedback circuit. A spike integrator stage captures the temporal activity of a set of neurons, and these activities are compared and digitized by the proposed WTA comparator stage. The proposed WTA designed in GF45RFSOI technology, exhibits self-excitation and global-inhibition properties, offers scalability, consumes 44% less power (140 nW ) and occupies a 40% lower area (166 μm2), compared to state-of-the-art. Gaurav R, Abhishek Kadam, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 6 |
| 2024 | A sub-100 nW Power, Compact CTDSM with a Band-To-Band Tunnelling Loop FilterabstractThis work presents a continuous-time delta-sigma modulator (CTDSM) deploying an experimentally demonstrated band-to-band-tunelling (BTBT) SOI MOSFET-based loop filter. With a compact, low-pass filter circuit and extremely low current in the BTBT regime, a loop filter implementation will provide optimality in terms of area and power performance. In literature for moderate-resolution CTDSMs, traditional loop filters are implemented with either fully passive, active, or hybrid integrators. These designs have a tight tradeoff in terms of area and power. The passive integrators have optimal power but suboptimal area, while the active integrators have optimal area and sub-optimal power consumption. The proposed work tries to break this tradeoff using BTBT regime loop filters. The CTDSM was designed in a GF45RFSOI technology and achieves a peak SNR/SNDR of 48.41 dB/47.94 dB for a 5 kHz bandwidth. The power consumption is 76.3 nW, with an area of 102.7 μm2— more than 100x area reduction over previous state-of-the-art moderate-precision CTDSM designs. This makes the proposed CTDSM extremely compact and power-efficient compared to traditional state-of-the-art moderate-resolution DSMs. Atharva Raut, Abhishek Kadam, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 6 |
| 2023 | Robustness to Variability and Asymmetry of In-Memory On-Chip Training
Rohit K. Vartak, Vivek Saraswat, Udayan Ganguly |
ICANN (9) | 3 |
| 2023 | MAdapter: A Multimodal Adapter for Liquid State Machines configures the Input Layer for the same Reservoir to enable Vision and Speech ClassificationabstractThe human cortex is capable of multi-modal processing. Then, assuming that the biological model of the brain cortex is captured by Liquid State Machines (LSMs), it should be endowed with a certain universality and efficiency that enables the same reservoir to extract features from different modalities like image, video and speech. However, researchers working with single applications have tuned LSMs specifically for their target datasets and modalities - possibly to account for the vast difference in spiking statistics produced by each unique dataset. In this paper, we explore a strategy to define the input layer hyperparameters for the same Reservoir to exploit the generality and efficiency of a Reservoir Network towards different modalities of data. We show that a strategy to set the input layer hyperparameters such that the Reservoir activity (average spike rate) is tuned to the best performance for one modality/dataset works for other modalities as long as the input layer hyperparameters are tuned to ensure similar levels of input stimulus to the reservoir. We also propose to use time partitions of the output spike trains coming from the Reservoir to improve the classification performance for time-varying datasets. We attained test classification accuracy of 86.74% on Fashion-MNIST, 95.86% on Neuromorphic-MNIST and 86.2% on the full TI-46 speech classification dataset using input parameters transferred from MNIST optimization on the same 1000-neuron Reservoir. These are equivalent or superior to state of the art in LSM classification performance. Using time partitions further pushes up the test classification accuracy of Neuromorphic-MNIST to 98% and of TI-46 speech classification dataset to 92% which approaches Artificial Neural Network (ANN) and Convolutional Neural Network (CNN) levels of performance respectively. This work not only simplifies the task of identifying optimal parameters for the functioning of an LSM across multiple input modalities, but also unlocks the generalizing power of using the same Reservoir network for multiple tasks Anmol Biswas, Nivedya S. Nambiar, Kushal Kejriwal, Udayan Ganguly |
IJCNN | 4 |
| 2023 | Real-world Performance Estimation of Liquid State Machines for Spoken Digit ClassificationabstractLiquid State Machine (LSM) is a brain-inspired neural network architecture for solving temporal classification problems like speech recognition. The simple structure of LSM with a reservoir and single-layer classifier is attractive from a hardware implementation perspective. When the LSM is considered for low-power hardware implementation in real-world command word recognition tasks, challenges like nonidealities in sensor filter response and ambient noise become critical concerns. In this work, we evaluate the performance of LSM based on two aspects (1) ambient noise and (2) sensor/preprocessing circuit nonidealities. For Ambient noise, we use additive white gaussian noise (AWGN) and ambient noise using the iNoise Indian Noise dataset that covers various natural indoor, outdoor, and travel-related environmental sounds. To understand the impact of input hardware nonidealities, we analyzed the impact of the audio preprocessing filter's quality factor, order, center frequency variations, and output nonlinearity on LSM performance. We use the spoken digits classification in the TI-46 dataset. This paper's findings present design guidelines for the system designers intending to use liquid-state machines for speech classification tasks. In terms of filter design, first, there is a broad Q, order space for filter design where performance is high. We use the hardware-friendly parallel 4th order Butterworth bandpass filter model to provide a baseline 98% accuracy in speech classification tasks. Second, the performance of LSM degrades proportionally to the variation in the center frequency of the bandpass filters in the filter bank. Third, nonlinearity with the third-order harmonic of 50 dBc can be tolerated. Regarding ambient noise, our study shows that a 40 dB SNR for AWGN is sufficient for ideal performance. Second, the best case of “home” noise leads to a performance of 91.4%. Outdoor and travel noise reduce the classification performance to 78.8% and 62.4%, respectively. However, ideal performance is recovered if the signal to noise ratio (SNR) is increased, particularly by 10 dB in indoor conditions and 30 dB in outdoor conditions. Thus, our study presents an engineering evaluation for real-world spoken digit recognition using LSMs. Abhishek Kadam, Anmol Biswas, Vivek Saraswat, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
IJCNN | 7 |
| 2023 | E-STDP: A Spatio-Temporally Local Unsupervised Learning Rule for Sparse Coded Spiking Convolutional AutoencodersabstractSparse coding algorithms can be used to train a network to perform compression and image reconstruction similar to autoencoders. These algorithms have allowed unsupervised learning of biologically plausible image filters in Spiking Neural Networks (SNNs) that can be used for image classification tasks. These methods have the advantage of being more power efficient due to their sparsity and spiking nature. However, in the absence of a temporally local learning rule, the training step has to be done offline in a separate conventional computing framework. We propose a spike-based temporally local update rule that implements sparse coding in SNNs. We trained a Convolutional Neural Network (CNN) which achieves a test accuracy of 98.35% on the MNIST dataset. Furthermore, it improves the power efficiency by representing an MNIST digit by around 10,000 spikes which is 4x more efficient representation, in terms of bits required, compared to a standard non-spiking CNN. Since the updates are both temporally and spatially local, it also performs the required learning within a small multiple of 10,000 operations per sample compared to around 1,870,000 operations per sample required to train a typical non-spiking CNN having a similar structure. Such a spatio-temporally local learning rule opens the way for implementation on neuromorphic hardware without the systemic and operational inefficiencies of global information broadcast. Vineet Kotariya, Anmol Biswas, Udayan Ganguly |
IJCNN | 3 |
| 2023 | Optimizing Throughput and Latency of Static 5G Multicast Networks using Boltzmann MachinesabstractThe 5G networks transmit data highly directionally and at very high rates. As a result, the data is not broadcast to all the receivers in the network simultaneously and there exist time delays in receivers servicing. The sender's goal is to find an optimal beam divergence angle for the antenna such that the data multicast rate is maximised while minimising alignment delay. There is a need for finding optimal receiver servicing solutions efficiently with high frequency especially for edge devices. Here we model the above problem as a constrained optimization problem and solve it using a Boltzmann Machine. The corresponding objective function has two terms: (i) the data transmission delay and (ii) the alignment delay. In addition to this, a penalty is imposed on the sender to ensure that all receivers receive the data exactly once and all the sets of receivers are serviced in a geometrically meaningful manner. The objective function is mapped to the functional form of the energy function of a Boltzmann Machine. The output comprises a prescription of which sets of receivers are to be serviced and an optimum value of beam divergence angle. This work captures the output visualisation, and the effect of relevant constraints appropriately for the first time. Efficient hardware-accelerated implementations of Boltzmann Machines using in-memory computing principles and emerging memories like memristors greatly increase the significance of such mappings. Such mappings of latency minimization problems are expected to be crucial for developing 5G mobile multicast networks in an IoT (Internet-of- Things) setting. Vadlamani Madhav, Vivek Saraswat, Udayan Ganguly |
IJCNN | 3 |
| 2023 | ANN Inference enabled by Variability Mitigation using 2T-1R Bit Cell-based Design Space AnalysisabstractResistive RAM (RRAM) devices are compact and easy to fabricate with electrical inputs-based switching. Conductive-Bridge RRAM (CBRAM) is being developed to meet retention, endurance and reliability specifications by GlobalFoundries for typical 1-bit per cell digital storage. However, filamentary growth and rupture produce variability, and the resistance states can often span multiple orders. We explore whether digital-storage-focused CBRAM can support analog current readout-based Multiply-and-Accumulate operations in artificial neural network (ANN) applications. We explore the 2T-1R bit cell to tune the mean HRS/LRS ratio and to control the variability in HRS and LRS readouts. We use experimental CBRAM data and GlobalFoundries' 22FDX platform and demonstrate > 2 × reduction in HRS and LRS logscale variability and > 10 × higher HRS/LRS ratio for the 2T-1R bit cell. The strategy is successfully tested for two datasets – the simpler MNIST and the more complex FMNIST using system-level modeling of non-idealities like weight quantization, HRS/LRS ratio, and variability in the readout of each bit-cell. Such bit-cell design principles have general utility in exploiting variability-prone characteristics of emerging memories for excellent application-level performance. Shreyas Deshmukh, Vivek Saraswat, Venkatesh Gopinath, Rajesh Nair, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 7 |
| 2023 | Enhanced regularization for on-chip training using analog and temporary memory weights
Raghav Singhal, Vivek Saraswat, Shreyas Deshmukh, Sreenivas Subramoney, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
Neural Networks | 7 |
| 2022 | Liquid State Machine on Loihi: Memory Metric for Performance Prediction
Rajat Patel, Vivek Saraswat, Udayan Ganguly |
ICANN (3) | 3 |
| 2022 | Spiking-GAN: A Spiking Generative Adversarial Network Using Time-To-First-Spike CodingabstractSpiking Neural Networks (SNNs) have shown great potential in solving deep learning problems in an energy-efficient manner. However, they are still limited to simple classification tasks. In this paper, we propose Spiking-GAN, the first spike-based Generative Adversarial Network (GAN). It employs a kind of temporal coding scheme called time-to-first-spike coding. We train it using approximate backpropagation in the temporal domain. We use simple integrate-and-fire (IF) neurons with very high refractory period for our network which ensures a maximum of one spike per neuron. This makes the model much sparser than a spike rate-based system. Our modified temporal loss function called ‘Aggressive TTFS’ improves the inference time of the network by over 33% and reduces the number of spikes in the network by more than 11% compared to previous works. Our experiments show that on training the network on the MNIST dataset using this approach, we can generate high quality samples with 57x lower energy consumption compared to ANN-based GANs. Thereby demonstrating the potential of this framework for solving such problems in the spiking domain. Vineet Kotariya, Udayan Ganguly |
IJCNN | 2 |
| 2022 | Quantum Tunneling Based Ultra-Compact and Energy Efficient Spiking Neuron Enables Hardware SNNabstractLow-power and low-area neurons are essential for hardware implementation of large-scale SNNs. Various novel-physics-based leaky-integrate-and-fire (LIF) neuron architectures have been proposed with low power and area, but are not compatible with CMOS technology to enable brain scale implementation of SNN. In this paper, for the first time, we demonstrate hardware implementation of recurrent SNN using proposed low-power, low-area, and low-leakage band-to-band-tunneling (BTBT) based neurons. A low-power thresholding circuit is proposed. We further propose a predistortion technique to linearize a nonlinear neuron without any area and power overhead. We establish the equivalence of the proposed neuron with the ideal LIF neuron to demonstrate its versatility. The tunneling regime enables a high input impedance in the BTBT neurons (few$\text{G}\Omega$) to enable a voltage input without loading the synaptic array. To verify the effect of the proposed neuron, a 36-neuron recurrent SNN is fabricated in GF-45nm PDSOI technology. We achieved 5000x lower energy-per-spike at a similar area and 10x lower standby power at a similar area and energy-per-spike. Such overall performance improvement enables brain scale computing. Ajay Kumar Singh, Vivek Saraswat, Maryam Shojaei Baghini, Udayan Ganguly |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Algorithm for 3D-Chemotaxis Using Spiking Neural Network
Jayesh Choudhary, Vivek Saraswat, Udayan Ganguly |
ICANN (5) | 3 |
| 2021 | Simplified Klinokinesis using Spiking Neural Networks for Resource-Constrained Navigation on the N euromorphic Processor LoihiabstractC. elegans shows chemotaxis using klinokinesis where the worm senses the concentration based on a single concentration sensor to compute the concentration gradient to perform foraging through gradient ascent/descent towards the target concentration followed by contour tracking. The biomimetic implementation requires complex neurons with multiple ion channel dynamics as well as interneurons for control. While this is a key capability of autonomous robots, its implementation on energy-efficient neuromorphic hardware like Intel's Loihi requires adaptation of the network to hardware-specific constraints, which has not been achieved. In this paper, we demonstrate the adaptation of chemotaxis based on klinokinesis to Loihi by implementing necessary neuronal dynamics with only LIF neurons as well as a complete spike-based implementation of all functions e.g. Heaviside function and subtractions. Our results show that Loihi implementation is equivalent to the software counterpart on Python in terms of performance - both during foraging and contour tracking. The Loihi results are also resilient in noisy environments. Thus, we demonstrate a successful adaptation of chemotaxis on Loihi - which can now be combined with the rich array of SNN blocks for SNN based complex robotic control. Apoorv Kishore, Vivek Saraswat, Udayan Ganguly |
IJCNN | 3 |
| 2021 | Hardware-Friendly Synaptic Orders and Timescales in Liquid State Machines for Speech ClassificationabstractLiquid State Machines are brain inspired spiking neural networks (SNNs) with random reservoir connectivity and bio-mimetic neuronal and synaptic models. Reservoir computing networks are proposed as an alternative to deep neural networks to solve temporal classification problems. Previous studies suggest 2ndorder (double exponential) synaptic waveform to be crucial for achieving high accuracy for TI-46 spoken digits recognition. The proposal of long-time range (ms) bio-mimetic synaptic waveforms is a challenge to compact and power efficient neuromorphic hardware. In this work, we analyze the role of synaptic orders namely:$\delta$(high output for single time step), 0th(rectangular with a finite pulse width), 1st(exponential fall) and 2ndorder (exponential rise and fall) and synaptic timescales on the reservoir output response and on the TI-46 spoken digits classification accuracy under a more comprehensive parameter sweep. We find the optimal operating point to be correlated to an optimal range of spiking activity in the reservoir. Further, the proposed 0thorder synapses perform at par with the biologically plausible 2ndorder synapses. This is substantial relaxation for circuit designers as synapses are the most abundant components in an in-memory implementation for SNNs. The circuit benefits for both analog and mixed-signal realizations of 0thorder synapse are highlighted demonstrating 2–3 orders of savings in area and power consumptions by eliminating Op-Amps and Digital to Analog Converter circuits. This has major implications on a complete neural network implementation with focus on peripheral limitations and algorithmic simplifications to overcome them. Vivek Saraswat, Ajinkya Gorad, Anand Naik, Aakash Patil, Udayan Ganguly |
IJCNN | 5 |
| 2020 | Adaptive Chemotaxis for Improved Contour Tracking Using Spiking Neural Networks
Shashwat Shukla, Rohan Pathak, Vivek Saraswat, Udayan Ganguly |
ICANN (2) | 4 |
| 2020 | Software-Level Accuracy Using Stochastic Computing With Charge-Trap-Flash Based Weight MatrixabstractThe in-memory computing paradigm with emerging memory devices has been recently shown to be a promising way to accelerate deep learning. Resistive processing unit (RPU) has been proposed to enable the vector-vector outer product in a crossbar array using a stochastic train of identical pulses to enable one-shot weight update, promising intense speed-up in matrix multiplication operations, which form the bulk of training neural networks. However, the performance of the system suffers if the device does not satisfy the condition of linear conductance change over around 1,000 conductance levels. This is a challenge for nanoscale memories. Recently, Charge Trap Flash (CTF) memory was shown to have a large number of levels before saturation, but variable non-linearity. In this paper, we explore the trade-off between the range of conductance change and linearity. We show, through simulations, that at an optimum choice of the range, our system performs nearly as well as the models trained using exact floating point operations, with less than 1% reduction in the performance. Our system reaches an accuracy of 97.9% on MNIST dataset, 89.1% and 70.5% accuracy on CIFAR-10 and CIFAR-100 datasets (using pre-extracted features). We also show its use in reinforcement learning, where it is used for value function approximation in Q-Learning, and learns to complete an episode the mountain car control problem in around 146 steps. Benchmarked to state-of-the-art, the CTF based RPU shows best in class performance to enable software equivalent performance. Varun Bhatt, Shrivastava Shalini, Tanmay Chavan, Udayan Ganguly |
IJCNN | 4 |
| 2020 | n-Oscillator Neural Network based Efficient Cost Function for n-city Traveling Salesman ProblemabstractNeural Networks have long been a mainstream technique to solve optimization problems. A classic example is the Travelling Salesman Problem (TSP) which NP-hard. Using a Hopfield-Tank representation, an n-city problem is mapped to a cost function of n2interacting neural units. Stochastic gradient descent helps achieve the global minima. Due to the nature of the TSP problem, the cost function has to penalize invalid sub-routes (non-Hamiltonian cycles) and minimize the travel cost simultaneously. In addition, there is a starting point and travel direction associated `2n' degeneracy. Previously, a cellular neuronal approach was proposed where the neural units were replaced with oscillators. The phase relations determined the output solution. Multiphase clusters of these oscillators solved the degeneracy issue. This paper proposes an n-oscillator cost function for an n -city TSP. Since a group of single frequency oscillator phases are naturally ordered and circular in a system, the proposed method exploits the true potential of oscillator nodes. The sub-routes and degeneracy are eliminated by design in addition to massively increasing the scaling potential (n vs. n2). It was also found that the proposed n- mapping can converge to the optimum tour much faster (about 100 times for a 5-city problem) than for n2mapping. Our approach projects hardware efficiency in terms of area footprint, computation time and energy. With coupled single device-based compact nanoscale oscillator systems becoming increasingly viable in hardware, efficient cost function mappings of hard problems using oscillator phases, as shown here, is critical to solving large graphical optimization problems. Shruti Landge, Vivek Saraswat, Srisht Fateh Singh, Udayan Ganguly |
IJCNN | 4 |
| 2020 | Circuit Cost Reduction for Online STDP using NIPIN Selector as Timekeeping Device in RRAM SynapseabstractOn-chip implementation of spike-time dependent plasticity in spiking neural networks using RRAM synapses requires pulse shaping circuits (PSC) to drive RRAMs. PSCs convert the temporal separation between pre and post neuron spikes to appropriate voltages that get applied across the synapse. The speculation of PSCs consuming the majority of circuit resources in the neuron circuits calls for methods simplifying the PSC. A recently demonstrated NIPIN timekeeping device based selector facilitates this, showing learning with square pulses using its inherent hole storage physics. However, a quantitative advantage achieved by utilizing a timekeeping device to evaluate its necessity is unavailable in the literature. Also, a model is required to carry out large scale circuit simulations for crossbar arrays using this device as selector. In this work, we design and compare the PSCs for different selector devices proposed in the literature to show 133× reduction in energy per spike and 8× reduction in the area of neuron circuit using NIPIN as the selector device compared to previously shown diode selector. We also present an experimentally calibrated model for the device for future explorations. Our results show that the small fraction energy and area occupied by the leaky-integrate and fire part of the circuit makes optimization of PSCs a priority. Thus, our work highlights the importance of mimicking biology by the use of simple spikes from neurons and performing time-keeping at the synapse in implementations of learning circuits. Ashwin Sanjay Lele, Anand Naik, Lakshya Bandhu, Bhaskar Das, Udayan Ganguly |
ISCAS | 5 |
| 2019 | Predicting Performance using Approximate State Space Model for Liquid State MachinesabstractLiquid State Machine (LSM) is a brain-inspired architecture used for solving problems like speech recognition and time series prediction. LSM comprises of a randomly connected recurrent network of spiking neurons. This network propagates the non-linear neuronal and synaptic dynamics. Maass et al. have argued that the non-linear dynamics of LSM is essential for its performance as a universal computer. Lyapunov exponent (μ), used to characterize the non-linearity of the network, correlates well with LSM performance. We propose a complementary approach of approximating the LSM dynamics with a linear state space representation. The spike rates from this model are well correlated to the spike rates from LSM. Such equivalence allows the extraction of a memory metric (τM) from the state transition matrix. τMdisplays high correlation with performance. Further, high τMsystems require fewer epochs to achieve a given accuracy. Being computationally cheap (1800× time efficient compared to LSM), the τMmetric enables exploration of the vast parameter design space. We observe that the performance correlation of the τMsurpasses that of Lyapunov exponent (μ), (2 - 4× improvement) in the high-performance regime over multiple datasets. In fact, while μ increases monotonically with network activity, the performance reaches a maxima at a specific activity described in literature as the edge of chaos. On the other hand, τMremains correlated with LSM performance. Hence, τMcaptures the useful memory of network activity that enables LSM performance. It also enables rapid design space exploration and fine-tuning of LSM parameters for high performance. Ajinkya Gorad, Vivek Saraswat, Udayan Ganguly |
IJCNN | 3 |
| 2018 | Sparsity Enables Data and Energy Efficient Spiking Convolutional Neural Networks
Varun Bhatt, Udayan Ganguly |
ICANN (1) | 2 |
| 2018 | Design of Spiking Rate Coded Logic Gates for C. elegans Inspired Contour Tracking
Shashwat Shukla, Sangya Dutta, Udayan Ganguly |
ICANN (1) | 3 |
| 2018 | A case for multiple and parallel RRAMs as synaptic model for training SNNsabstractTo enable a dense integration of model synapses in a spiking neural networks (SNN) hardware, various nanoscale devices are being considered. Such devices, besides exhibiting spike-timing dependent plasticity (STDP), need to be highly scalable, have a large endurance and require low energy for transitioning between states. In this work, first, we introduce and empirically determine two new specifications for a resistive random-access memory (RRAM) based synapse: number of conductance levels per synapse and learning-rate. To the best of our knowledge, there are no RRAMs that meet the latter specification. As a solution, we propose the use of multiple RRAMs in parallel within a synapse. While synaptic reading, all RRAMs are simultaneously read and for each synaptic conductance-change event, the mechanism for conductance STDP is initiated on only one RRAM, randomly picked from the set. Second, to validate our solution, we experimentally demonstrate STDP of conductance of a Pr0.7Ca0.3MnO3(PCMO)-RRAM and then show that due to a large learning-rate, a single PCMO-RRAM fails to model a synapse in the training of an SNN. As anticipated, network training improved as more PCMO-RRAMs were added to the synapse. Fourth, we discuss circuit-requirements for implementing such a scheme, to conclude that the requirements are within bounds. Thus, our work presents specifications for synaptic devices in trainable SNNs, indicates the shortcomings of state-of-art synaptic contenders, and provides a solution to extrinsically meet the specifications and discusses the peripheral circuitry that implements the solution. Sidharth Prasad, Sandip Lashkare, Udayan Ganguly |
IJCNN | 4 |
| 2018 | Stochastic learning in deep neural networks based on nanoscale PCMO device characteristics
Anakha V. Babu, Sandip Lashkare, Udayan Ganguly, Bipin Rajendran |
Neurocomputing | 3 |
| 2017 | A software-equivalent SNN hardware using RRAM-array for asynchronous real-time learningabstractSpiking Neural Network (SNN) naturally inspires hardware implementation as it is based on biology. For learning, spike time dependent plasticity (STDP) may be implemented using an energy efficient waveform superposition on memristor based synapse. However, system level implementation has three challenges. First, a classic dilemma is that recognition requires current reading for short voltage-spikes which is disturbed by large voltage-waveforms that are simultaneously applied on the same memristor for real-time learning i.e. the simultaneous read-write dilemma. Second, the hardware needs to exactly replicate software implementation for easy adaptation of algorithm to hardware. Third, the devices used in hardware simulations must be realistic. In this paper, we present an approach to address the above concerns. First, the learning and recognition occurs in separate arrays simultaneously in real-time, asynchronously — avoiding non-biomimetic clocking based complex signal management. Second, we show that the hardware emulates software at every stage by comparison of SPICE (circuit-simulator) with MATLAB® (mathematical SNN algorithm implementation in software) implementations. As an example, the hardware shows 97.5% accuracy in classification which is equivalent to software for a Fisher's Iris dataset. Third, the STDP is implemented using a model of synaptic device implemented using HfO2memristor. We show that an increasingly realistic memristor model slightly reduces the hardware performance (85%), which highlights the need to engineer RRAM characteristics specifically for SNN. Udayan Ganguly |
IJCNN | 3 |