EDBT 2026 Demo / reviewers in the wild / expert
Vivek Saraswat
dblp:234/7909
· DBLP profile ↗
14ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0001-9191-1632ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analog and Temporary On-chip Memory for ANN Training and InferenceabstractOn-chip training at the edge becomes a primary requisite for real-time and security-sensitive artificial neural network (ANN) applications. In-memory computation (IMC) techniques have been proposed to facilitate data-intensive computational operations in ANNs. IMC-based multiply-accumulate (MAC) accelerates ANN training but suffers from significant communication overhead between the MAC engine and the off-chip storage for the intermediate data. This article proposes an analog temporary on-chip memory (ATOM) to store this intermediate data during ANN training. The ANN training architecture with the proposed ATOM has two significant advantages. First, the energy required to store intermediate data is scaled down by \(\sim\) 40 \(\times\) due to the on-chip and analog nature of the memory. Second, the proposed architecture avoids power and area-consuming analog-to-digital converters (ADCs) between neural network stages. The ATOM cell measurements are carried out from 20 fabricated chips, and the impact of ATOM characteristics on ANN system performance accuracy is analyzed. This article shows significant latency improvement of \(\sim\) 9 \(\times\) and area savings of \(\sim\) 5 \(\times\) for intermediate data storage compared to the on-chip SRAM during ANN training’s forward and backward pass operations. An improvement in the area and latency will be beneficial to instrument the area- and energy-efficient hardware system for on-chip ANN applications. Shreyas Deshmukh, Raghav Singhal, Shruti Landge, Vivek Saraswat, Anmol Biswas, Abhishek Kadam, Ajay Kumar Singh, Sreenivas Subramoney, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2023 | Robustness to Variability and Asymmetry of In-Memory On-Chip Training
Rohit K. Vartak, Vivek Saraswat, Udayan Ganguly |
ICANN (9) | 2 |
| 2023 | Real-world Performance Estimation of Liquid State Machines for Spoken Digit ClassificationabstractLiquid State Machine (LSM) is a brain-inspired neural network architecture for solving temporal classification problems like speech recognition. The simple structure of LSM with a reservoir and single-layer classifier is attractive from a hardware implementation perspective. When the LSM is considered for low-power hardware implementation in real-world command word recognition tasks, challenges like nonidealities in sensor filter response and ambient noise become critical concerns. In this work, we evaluate the performance of LSM based on two aspects (1) ambient noise and (2) sensor/preprocessing circuit nonidealities. For Ambient noise, we use additive white gaussian noise (AWGN) and ambient noise using the iNoise Indian Noise dataset that covers various natural indoor, outdoor, and travel-related environmental sounds. To understand the impact of input hardware nonidealities, we analyzed the impact of the audio preprocessing filter's quality factor, order, center frequency variations, and output nonlinearity on LSM performance. We use the spoken digits classification in the TI-46 dataset. This paper's findings present design guidelines for the system designers intending to use liquid-state machines for speech classification tasks. In terms of filter design, first, there is a broad Q, order space for filter design where performance is high. We use the hardware-friendly parallel 4th order Butterworth bandpass filter model to provide a baseline 98% accuracy in speech classification tasks. Second, the performance of LSM degrades proportionally to the variation in the center frequency of the bandpass filters in the filter bank. Third, nonlinearity with the third-order harmonic of 50 dBc can be tolerated. Regarding ambient noise, our study shows that a 40 dB SNR for AWGN is sufficient for ideal performance. Second, the best case of “home” noise leads to a performance of 91.4%. Outdoor and travel noise reduce the classification performance to 78.8% and 62.4%, respectively. However, ideal performance is recovered if the signal to noise ratio (SNR) is increased, particularly by 10 dB in indoor conditions and 30 dB in outdoor conditions. Thus, our study presents an engineering evaluation for real-world spoken digit recognition using LSMs. Abhishek Kadam, Anmol Biswas, Vivek Saraswat, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
IJCNN | 3 |
| 2023 | Optimizing Throughput and Latency of Static 5G Multicast Networks using Boltzmann MachinesabstractThe 5G networks transmit data highly directionally and at very high rates. As a result, the data is not broadcast to all the receivers in the network simultaneously and there exist time delays in receivers servicing. The sender's goal is to find an optimal beam divergence angle for the antenna such that the data multicast rate is maximised while minimising alignment delay. There is a need for finding optimal receiver servicing solutions efficiently with high frequency especially for edge devices. Here we model the above problem as a constrained optimization problem and solve it using a Boltzmann Machine. The corresponding objective function has two terms: (i) the data transmission delay and (ii) the alignment delay. In addition to this, a penalty is imposed on the sender to ensure that all receivers receive the data exactly once and all the sets of receivers are serviced in a geometrically meaningful manner. The objective function is mapped to the functional form of the energy function of a Boltzmann Machine. The output comprises a prescription of which sets of receivers are to be serviced and an optimum value of beam divergence angle. This work captures the output visualisation, and the effect of relevant constraints appropriately for the first time. Efficient hardware-accelerated implementations of Boltzmann Machines using in-memory computing principles and emerging memories like memristors greatly increase the significance of such mappings. Such mappings of latency minimization problems are expected to be crucial for developing 5G mobile multicast networks in an IoT (Internet-of- Things) setting. Vadlamani Madhav, Vivek Saraswat, Udayan Ganguly |
IJCNN | 2 |
| 2023 | ANN Inference enabled by Variability Mitigation using 2T-1R Bit Cell-based Design Space AnalysisabstractResistive RAM (RRAM) devices are compact and easy to fabricate with electrical inputs-based switching. Conductive-Bridge RRAM (CBRAM) is being developed to meet retention, endurance and reliability specifications by GlobalFoundries for typical 1-bit per cell digital storage. However, filamentary growth and rupture produce variability, and the resistance states can often span multiple orders. We explore whether digital-storage-focused CBRAM can support analog current readout-based Multiply-and-Accumulate operations in artificial neural network (ANN) applications. We explore the 2T-1R bit cell to tune the mean HRS/LRS ratio and to control the variability in HRS and LRS readouts. We use experimental CBRAM data and GlobalFoundries' 22FDX platform and demonstrate > 2 × reduction in HRS and LRS logscale variability and > 10 × higher HRS/LRS ratio for the 2T-1R bit cell. The strategy is successfully tested for two datasets – the simpler MNIST and the more complex FMNIST using system-level modeling of non-idealities like weight quantization, HRS/LRS ratio, and variability in the readout of each bit-cell. Such bit-cell design principles have general utility in exploiting variability-prone characteristics of emerging memories for excellent application-level performance. Shreyas Deshmukh, Vivek Saraswat, Venkatesh Gopinath, Rajesh Nair, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 2 |
| 2023 | Enhanced regularization for on-chip training using analog and temporary memory weights
Raghav Singhal, Vivek Saraswat, Shreyas Deshmukh, Sreenivas Subramoney, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
Neural Networks | 2 |
| 2022 | Liquid State Machine on Loihi: Memory Metric for Performance Prediction
Rajat Patel, Vivek Saraswat, Udayan Ganguly |
ICANN (3) | 2 |
| 2022 | Quantum Tunneling Based Ultra-Compact and Energy Efficient Spiking Neuron Enables Hardware SNNabstractLow-power and low-area neurons are essential for hardware implementation of large-scale SNNs. Various novel-physics-based leaky-integrate-and-fire (LIF) neuron architectures have been proposed with low power and area, but are not compatible with CMOS technology to enable brain scale implementation of SNN. In this paper, for the first time, we demonstrate hardware implementation of recurrent SNN using proposed low-power, low-area, and low-leakage band-to-band-tunneling (BTBT) based neurons. A low-power thresholding circuit is proposed. We further propose a predistortion technique to linearize a nonlinear neuron without any area and power overhead. We establish the equivalence of the proposed neuron with the ideal LIF neuron to demonstrate its versatility. The tunneling regime enables a high input impedance in the BTBT neurons (few$\text{G}\Omega$) to enable a voltage input without loading the synaptic array. To verify the effect of the proposed neuron, a 36-neuron recurrent SNN is fabricated in GF-45nm PDSOI technology. We achieved 5000x lower energy-per-spike at a similar area and 10x lower standby power at a similar area and energy-per-spike. Such overall performance improvement enables brain scale computing. Ajay Kumar Singh, Vivek Saraswat, Maryam Shojaei Baghini, Udayan Ganguly |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Algorithm for 3D-Chemotaxis Using Spiking Neural Network
Jayesh Choudhary, Vivek Saraswat, Udayan Ganguly |
ICANN (5) | 2 |
| 2021 | Simplified Klinokinesis using Spiking Neural Networks for Resource-Constrained Navigation on the N euromorphic Processor LoihiabstractC. elegans shows chemotaxis using klinokinesis where the worm senses the concentration based on a single concentration sensor to compute the concentration gradient to perform foraging through gradient ascent/descent towards the target concentration followed by contour tracking. The biomimetic implementation requires complex neurons with multiple ion channel dynamics as well as interneurons for control. While this is a key capability of autonomous robots, its implementation on energy-efficient neuromorphic hardware like Intel's Loihi requires adaptation of the network to hardware-specific constraints, which has not been achieved. In this paper, we demonstrate the adaptation of chemotaxis based on klinokinesis to Loihi by implementing necessary neuronal dynamics with only LIF neurons as well as a complete spike-based implementation of all functions e.g. Heaviside function and subtractions. Our results show that Loihi implementation is equivalent to the software counterpart on Python in terms of performance - both during foraging and contour tracking. The Loihi results are also resilient in noisy environments. Thus, we demonstrate a successful adaptation of chemotaxis on Loihi - which can now be combined with the rich array of SNN blocks for SNN based complex robotic control. Apoorv Kishore, Vivek Saraswat, Udayan Ganguly |
IJCNN | 2 |
| 2021 | Hardware-Friendly Synaptic Orders and Timescales in Liquid State Machines for Speech ClassificationabstractLiquid State Machines are brain inspired spiking neural networks (SNNs) with random reservoir connectivity and bio-mimetic neuronal and synaptic models. Reservoir computing networks are proposed as an alternative to deep neural networks to solve temporal classification problems. Previous studies suggest 2ndorder (double exponential) synaptic waveform to be crucial for achieving high accuracy for TI-46 spoken digits recognition. The proposal of long-time range (ms) bio-mimetic synaptic waveforms is a challenge to compact and power efficient neuromorphic hardware. In this work, we analyze the role of synaptic orders namely:$\delta$(high output for single time step), 0th(rectangular with a finite pulse width), 1st(exponential fall) and 2ndorder (exponential rise and fall) and synaptic timescales on the reservoir output response and on the TI-46 spoken digits classification accuracy under a more comprehensive parameter sweep. We find the optimal operating point to be correlated to an optimal range of spiking activity in the reservoir. Further, the proposed 0thorder synapses perform at par with the biologically plausible 2ndorder synapses. This is substantial relaxation for circuit designers as synapses are the most abundant components in an in-memory implementation for SNNs. The circuit benefits for both analog and mixed-signal realizations of 0thorder synapse are highlighted demonstrating 2–3 orders of savings in area and power consumptions by eliminating Op-Amps and Digital to Analog Converter circuits. This has major implications on a complete neural network implementation with focus on peripheral limitations and algorithmic simplifications to overcome them. Vivek Saraswat, Ajinkya Gorad, Anand Naik, Aakash Patil, Udayan Ganguly |
IJCNN | 1 |
| 2020 | Adaptive Chemotaxis for Improved Contour Tracking Using Spiking Neural Networks
Shashwat Shukla, Rohan Pathak, Vivek Saraswat, Udayan Ganguly |
ICANN (2) | 3 |
| 2020 | n-Oscillator Neural Network based Efficient Cost Function for n-city Traveling Salesman ProblemabstractNeural Networks have long been a mainstream technique to solve optimization problems. A classic example is the Travelling Salesman Problem (TSP) which NP-hard. Using a Hopfield-Tank representation, an n-city problem is mapped to a cost function of n2interacting neural units. Stochastic gradient descent helps achieve the global minima. Due to the nature of the TSP problem, the cost function has to penalize invalid sub-routes (non-Hamiltonian cycles) and minimize the travel cost simultaneously. In addition, there is a starting point and travel direction associated `2n' degeneracy. Previously, a cellular neuronal approach was proposed where the neural units were replaced with oscillators. The phase relations determined the output solution. Multiphase clusters of these oscillators solved the degeneracy issue. This paper proposes an n-oscillator cost function for an n -city TSP. Since a group of single frequency oscillator phases are naturally ordered and circular in a system, the proposed method exploits the true potential of oscillator nodes. The sub-routes and degeneracy are eliminated by design in addition to massively increasing the scaling potential (n vs. n2). It was also found that the proposed n- mapping can converge to the optimum tour much faster (about 100 times for a 5-city problem) than for n2mapping. Our approach projects hardware efficiency in terms of area footprint, computation time and energy. With coupled single device-based compact nanoscale oscillator systems becoming increasingly viable in hardware, efficient cost function mappings of hard problems using oscillator phases, as shown here, is critical to solving large graphical optimization problems. Shruti Landge, Vivek Saraswat, Srisht Fateh Singh, Udayan Ganguly |
IJCNN | 2 |
| 2019 | Predicting Performance using Approximate State Space Model for Liquid State MachinesabstractLiquid State Machine (LSM) is a brain-inspired architecture used for solving problems like speech recognition and time series prediction. LSM comprises of a randomly connected recurrent network of spiking neurons. This network propagates the non-linear neuronal and synaptic dynamics. Maass et al. have argued that the non-linear dynamics of LSM is essential for its performance as a universal computer. Lyapunov exponent (μ), used to characterize the non-linearity of the network, correlates well with LSM performance. We propose a complementary approach of approximating the LSM dynamics with a linear state space representation. The spike rates from this model are well correlated to the spike rates from LSM. Such equivalence allows the extraction of a memory metric (τM) from the state transition matrix. τMdisplays high correlation with performance. Further, high τMsystems require fewer epochs to achieve a given accuracy. Being computationally cheap (1800× time efficient compared to LSM), the τMmetric enables exploration of the vast parameter design space. We observe that the performance correlation of the τMsurpasses that of Lyapunov exponent (μ), (2 - 4× improvement) in the high-performance regime over multiple datasets. In fact, while μ increases monotonically with network activity, the performance reaches a maxima at a specific activity described in literature as the edge of chaos. On the other hand, τMremains correlated with LSM performance. Hence, τMcaptures the useful memory of network activity that enables LSM performance. It also enables rapid design space exploration and fine-tuning of LSM parameters for high performance. Ajinkya Gorad, Vivek Saraswat, Udayan Ganguly |
IJCNN | 2 |