Gopalakrishnan Srinivasan

dblp:05/4913 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0003-2015-8545ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 SPOILER-GUARD: Gating Latency Effects of Memory Accesses through Randomized Dependency Prediction
Gayathri Subramanian, Girinath P, Nitya Ranganathan, V. Kamakoti 0001, Gopalakrishnan Srinivasan
DATE5
2025 QuAKE: Speeding up Model Inference Using Quick and Approximate Kernels for Exponential Non-Linearities
abstract
As machine learning gets deployed more and more widely, and model sizes continue to grow, improving computational efficiency during model inference has become a key challenge. In many commonly used model architectures, including Transformers, a significant portion of the inference computation is comprised of exponential non-linearities such as Softmax. In this work, we develop QuAKE, a collection of novel operators that leverage certain properties of IEEE-754 floating point representations to quickly approximate the exponential function without requiring specialized hardware, extra memory, or precomputation. We propose optimizations that enhance the efficiency of QuAKE in commonly used exponential non-linearities such as Softmax, GELU, and the Logistic function. Our benchmarks demonstrate substantial inference speed improvements between 10% and 35% on server CPUs, and 5% and 45% on embedded and mobile-scale CPUs for a variety of model architectures and sizes. Evaluations of model performance on standard datasets and tasks from various domains show that QuAKE operators are able to provide sizable speed benefits with little to no loss of performance on downstream tasks.
Sai Kiran Narayanaswami, Gopalakrishnan Srinivasan, Balaraman Ravindran
ISLPED2
2021 Complexity-aware Adaptive Training and Inference for Edge-Cloud Distributed AI Systems
abstract
The ubiquitous use of IoT and machine learning applications is creating large amounts of data that require accurate and real-time processing. Although edge-based smart data processing can be enabled by deploying pretrained models, the energy and memory constraints of edge devices necessitate distributed deep learning between the edge and the cloud for complex data. In this paper, we propose a distributed system to exploit both the edge and the cloud for training and inference. We propose a new architecture, MEANet, with a main block, an extension block, and an adaptive block for the edge. The inference process can terminate at either the main block, the extension block, or the cloud. MEANet is trained to categorize inputs into easy/hard/complex classes. The main block identifies instances of easy/hard classes and classifies easy classes with high confidence. Only data with high probabilities of belonging to hard classes would be sent to the extension block for prediction. Further, only if the neural network at the edge shows low confidence in the prediction, the instance is considered complex and sent to the cloud for further processing. The training technique lends to the majority of inference on edge devices while going to the cloud only for a small set of complex jobs. The performance of the proposed system is evaluated via extensive experiments using modified models of ResNets and MobileNetV2 on CIFAR-100 and ImageNet datasets. The results show that the proposed distributed model has improved accuracy and lower energy consumption compared to standard models, indicating its capacity to adapt.
Yinghan Long, Indranil Chakraborty, Gopalakrishnan Srinivasan, Kaushik Roy 0001
ICDCS3
2020 RMP-SNN: Residual Membrane Potential Neuron for Enabling Deeper High-Accuracy and Low-Latency Spiking Neural Network
abstract
Spiking Neural Networks (SNNs) have recently attracted significant research interest as the third generation of artificial neural networks that can enable low-power event-driven data analytics. The best performing SNNs for image recognition tasks are obtained by converting a trained Analog Neural Network (ANN), consisting of Rectified Linear Units (ReLU), to SNN composed of integrate-and-fire neurons with "proper" firing thresholds. The converted SNNs typically incur loss in accuracy compared to that provided by the original ANN and require sizable number of inference time-steps to achieve the best accuracy. We find that performance degradation in the converted SNN stems from using "hard reset" spiking neuron that is driven to fixed reset potential once its membrane potential exceeds the firing threshold, leading to information loss during SNN inference. We propose ANN-SNN conversion using "soft reset" spiking neuron model, referred to as Residual Membrane Potential (RMP) spiking neuron, which retains the "residual" membrane potential above threshold at the firing instants. We demonstrate near loss-less ANN-SNN conversion using RMP neurons for VGG-16, ResNet-20, and ResNet-34 SNNs on challenging datasets including CIFAR-10 (93.63% top-1), CIFAR-100 (70.93% top-1), and ImageNet (73.09% top-1 accuracy). Our results also show that RMP-SNN surpasses the best inference accuracy provided by the converted SNN with "hard reset" spiking neurons using 2-8 times fewer inference time-steps across network architectures and datasets.
Bing Han 0006, Gopalakrishnan Srinivasan, Kaushik Roy 0001
CVPR2
2020 Training Deep Spiking Neural Networks for Energy-Efficient Neuromorphic Computing
abstract
Spiking Neural Networks (SNNs), widely known as the third generation of neural networks, encode input information temporally using sparse spiking events, which can be harnessed to achieve higher computational efficiency for cognitive tasks. However, considering the rapid strides in accuracy enabled by state-of-the-art Analog Neural Networks (ANNs), SNN training algorithms are much less mature, leading to accuracy gap between SNNs and ANNs. In this paper, we propose different SNN training methodologies, varying in degrees of biofidelity, and evaluate their efficacy on complex image recognition datasets. First, we present biologically plausible Spike Timing Dependent Plasticity (STDP) based deterministic and stochastic algorithms for unsupervised representation learning in SNNs. Our analysis on the CIFAR-10 dataset indicates that STDP-based learning rules enable the convolutional layers to self-learn low-level input features using fewer training examples. However, STDP-based learning is limited in applicability to shallow SNNs (≤4 layers) while yielding considerably lower than state-of-the-art accuracy. In order to scale the SNNs deeper and improve the accuracy further, we propose conversion methodology to map off-the-shelf trained ANN to SNN for energy-efficient inference. We demonstrate 69.96% accuracy for VGG16-SNN on ImageNet. However, ANN-to-SNN conversion leads to high inference latency for achieving the best accuracy. In order to minimize the inference latency, we propose spike-based error backpropagation algorithm using differentiable approximation for the spiking neuron. Our preliminary experiments on CIFAR-10 show that spike-based error backpropagation effectively captures temporal statistics to reduce the inference latency by up to 8× compared to converted SNNs while yielding comparable accuracy
Gopalakrishnan Srinivasan, Chankyu Lee, Abhronil Sengupta, Priyadarshini Panda, Syed Shakib Sarwar, Kaushik Roy 0001
ICASSP1
2020 Enabling Deep Spiking Neural Networks with Hybrid Conversion and Spike Timing Dependent Backpropagation
Nitin Rathi, Gopalakrishnan Srinivasan, Priyadarshini Panda, Kaushik Roy 0001
ICLR2
2020 Enabling Homeostasis using Temporal Decay Mechanisms in Spiking CNNs Trained with Unsupervised Spike Timing Dependent Plasticity
abstract
Convolutional Neural Networks(CNNs) have become the work horse for image classification tasks. This success has driven the exploration of Spike Time Dependent Plasticity (STDP) learning rule applied to the convolutional architecture for complex datasets as opposed to the fully connected architecture. Inhibitory neurons and adaptive threshold are widely adopted methods of inducing homeostasis in fully connected spiking networks to aid the unsupervised learning process. These methods ensure that all neurons have approximately equal firing activity across time and that their receptive fields are different, generally referred to as homeostatic behavior. While the adaptive threshold is straightforward to implement in spiking CNNs, adding in-hibitory neurons is not suitable to the convolutional architecture due to its shared weight nature. In this work, we first show that adaptive threshold in isolation is weak in obtaining approximate equal firing activity across activation maps in a spiking CNN. Next, we develop weight and offset decay mechanisms that enable the desired behavior to complement the STDP learning rule and adaptive threshold. We empirically show that these decay mechanisms improve feature learning as compared to baseline STDP in terms of accuracy (up to 1.4%) as well as enhanced homeostatic behavior among activation maps (more than halving the standard deviation). We discuss the complementary behavior of the decay mechanisms as compared to the adaptive threshold in terms of the variance in the activity induced. Finally, we show that when the convolutional features are trained on a subset of classes using STDP with decay mechanisms, the features learned are transferable to the subset of classes that are unseen to the convolutional layers. Thus, the decay mechanisms not only encourage the network to learn better features corresponding to the task being trained for but learn common structure prevalent among the classes while encouraging contribution from all activation maps. We perform experiments and present our findings on the Extended MNIST (EMNIST) dataset.
Krishna Reddy Kesari, Priyadarshini Panda, Gopalakrishnan Srinivasan, Kaushik Roy 0001
IJCNN3
2020 Pruning Filters while Training for Efficiently Optimizing Deep Learning Networks
abstract
Deep Neural Networks are an important class of machine learning algorithms that have demonstrated state-of-the-art accuracy for different cognitive tasks like image and speech recognition. Modem deep networks have millions to billions of parameters, which leads to high memory and energy requirements during training as well as during inference on resource-constrained edge devices. Consequently, pruning techniques have been proposed that remove less significant weights in deep networks, thereby reducing their memory and computational requirements. Pruning is usually performed after training the original network, and is followed by further retraining to compensate for the accuracy loss incurred during pruning. The prune-and-retrain procedure is repeated iteratively until an optimum tradeoff between accuracy and efficiency is reached. However, such iterative retraining adds to the overall training complexity of the network. In this work, we propose a dynamic pruning-while-training procedure, wherein we prune filters of the convolutional layers of a deep neural network during training itself, thereby precluding the need for separate retraining. We evaluate our dynamic pruning-while-training approach with three different pre-existing pruning strategies, viz. mean activation-based pruning, random pruning, and L1 normalization-based pruning. Our results for VGG-16 trained on CIFAR10 shows that L1 normalization provides the best performance among all the techniques explored in this work with less than 1% drop in accuracy after pruning 80% of the filters compared to the original network. We further evaluated the L1 normalization based pruning mechanism on CIFAR100. Results indicate that pruning while training yields a compressed network with almost no accuracy loss after pruning 50% of the filters compared to the original network and ~5% loss for high pruning rates (> 80%). The proposed pruning methodology yields 41% reduction in the number of computations and memory accesses during training for CIFAR10, CIFAR100 and ImageNet compared to training with retraining for 10 epochs.
Sourjya Roy, Priyadarshini Panda, Gopalakrishnan Srinivasan, Anand Raghunathan
IJCNN3
2020 Revisiting Stochastic Computing in the Era of Nanoscale Nonvolatile Technologies
abstract
In this era of nanoscale technologies, the inherent characteristics of some nonvolatile devices, such as resistive random access memory (ReRAM), phase-change material (PCM), and spintronics, can emulate stochastic functionalities. Traditionally, these devices have been engineered to suppress the stochastic switching behavior as it poses reliability concerns for memory storage and logic applications. However, leveraging stochasticity in such devices led to a renewed interest in hardware-software codesign of stochastic algorithms since the CMOS-based implementations of stochastic algorithms involve cumbersome circuitry to generate “stochastic bits.” In this article, we consider two classes of problems: deep neural networks (DNNs) and combinatorial optimization. The rapidly growing demands of artificial intelligence (AI) have sparked an interest in energy-efficient implementations of large DNNs, with binary representations of synaptic weights and neuronal activities. Stochasticity plays an important role in leveraging the benefits of these binary representations, leading to model compression and optimization during training. In combinatorial optimization, such as graph coloring or traveling salesman problems, stochastic algorithms, such as the Ising computing model, have been shown to be effective. These problems require exhaustive computational procedures, and the Ising model uses a natural annealing agent to achieve near-optimal solutions in a reasonable timescale, without getting stuck in “local minima.” In this article, we present a broad review of stochastic computing utilizing the stochastic switching characteristics of devices based on nanoscale nonvolatile technologies. We show how to codesign of the devices and algorithms that can enable optimal solutions for both combinatorial problems and binary neural networks for local learning and inference. Directly mapping the nonvolatile device characteristics to the stochastic algorithms without the need for storing the bits in a separate memory leads to efficient use of hardware.
Amogh Agrawal, Indranil Chakraborty, Deboleena Roy, Utkarsh Saxena, Saima Sharmin, Minsuk Koo, Yong Shim, Gopalakrishnan Srinivasan, Chamika M. Liyanagedera, Abhronil Sengupta, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.8
2019 Neural Networks at the Edge
abstract
As neural networks gain importance with several successful applications of them, this paper raises the question of how they can be applied in the context of coalition operations. A key challenge in military coalition operations is that of energy and severe bandwidth constraints. We address this challenge by exploring the use of Deep Neural Networks (DNNs) and splitting them across multiple edge nodes. Further, we explore the idea of using spiking neural networks that can lower the energy consumption significantly. Preliminary results show that both these approaches can have significant impact on coalition operations.
Deboleena Roy, Gopalakrishnan Srinivasan, Priyadarshini Panda, Richard Tomsett, Nirmit Desai, Raghu K. Ganti, Kaushik Roy 0001
SMARTCOMP2
2018 STDP-based Unsupervised Feature Learning using Convolution-over-time in Spiking Neural Networks for Energy-Efficient Neuromorphic Computing
abstract
Brain-inspired learning models attempt to mimic the computations performed in the neurons and synapses constituting the human brain to achieve its efficiency in cognitive tasks. In this work, we propose Spike Timing Dependent Plasticity-based unsupervised feature learning using convolution-over-time in Spiking Neural Network (SNN). We use shared weight kernels that are convolved with the input patterns over time to encode representative input features, thereby improving the sparsity as well as the robustness of the learning model. We show that the Convolutional SNN self-learns several visual categories for object recognition with limited number of training patterns while yielding comparable classification accuracy relative to the fully connected SNN. Further, we quantify the energy benefits of the Convolutional SNN over fully connected SNN on neuromorphic hardware implementation.
Gopalakrishnan Srinivasan, Priyadarshini Panda, Kaushik Roy 0001
ACM J. Emerg. Technol. Comput. Syst.1
2017 Magnetic tunnel junction enabled all-spin stochastic spiking neural network
abstract
Biologically-inspired spiking neural networks (SNNs) have attracted significant research interest due to their inherent computational efficiency in performing classification and recognition tasks. The conventional CMOS-based implementations of large-scale SNNs are power intensive. This is a consequence of the fundamental mismatch between the technology used to realize the neurons and synapses, and the neuroscience mechanisms governing their operation, leading to area-expensive circuit designs. In this work, we present a three-terminal spintronic device, namely, the magnetic tunnel junction (MTJ)-heavy metal (HM) heterostructure that is inherently capable of emulating the neuronal and synaptic dynamics. We exploit the stochastic switching behavior of the MTJ in the presence of thermal noise to mimic the probabilistic spiking of cortical neurons, and the conditional change in the state of a binary synapse based on the pre- and post-synaptic spiking activity required for plasticity. We demonstrate the efficacy of a crossbar organization of our MTJ-HM based stochastic SNN in digit recognition using a comprehensive device-circuit-system simulation framework. The energy efficiency of the proposed system stems from the ultra-low switching energy of the MTJ-HM device, and the in-memory computation rendered possible by the localized arrangement of the computational units (neurons) and non-volatile synaptic memory in such crossbar architectures.
Gopalakrishnan Srinivasan, Abhronil Sengupta, Kaushik Roy 0001
DATE1
2017 EnsembleSNN: Distributed assistive STDP learning for energy-efficient recognition in spiking neural networks
abstract
We present an ensemble approach for implementing Spiking Neural Networks (SNNs) with on-line unsupervised learning, well-suited for robust and energy-efficient design of neuromorphic computing systems for pattern recognition tasks. Inspired from the collective neuronal activity observed in the visual cortex, the proposed EnsembleSNN architecture involves multiple simple SNNs or ensembles acting in parallel on different aspects of the input. This in turn reduces the training complexity due to the decreased connectivity obtained from decomposing the input across different ensembles. During inference, a collective decision from all ensembles of the EnsembleSNN is considered to obtain the final prediction. We add predictive connections across different ensembles that enables individual ensembles to learn some statistics about the remaining portions of the input image that further enhances the collective decision making of the proposed architecture. We evaluate our approach on the MNIST dataset for different configurations of EnsembleSNN. Our experiments demonstrate upto 2.8x improvement in efficiency, while yielding better (~2.5%) accuracy than the optimized baseline network, and even higher improvements of upto 3.7x for minimal accuracy degradation (~3.2%).
Priyadarshini Panda, Gopalakrishnan Srinivasan, Kaushik Roy 0001
IJCNN2
2017 Spike timing dependent plasticity based enhanced self-learning for efficient pattern recognition in spiking neural networks
abstract
Spike Timing Dependent Plasticity (STDP), wherein synaptic weights are modified based on the temporal correlation between a pair of pre- and post-synaptic (post-neuronal) spikes, is widely used to implement unsupervised learning in Spiking Neural Networks (SNNs). In general, STDP-based learning models disregard the information embedded in post-neuronal spiking frequency. We observe that updating the synaptic weights at the instants of every post-neuronal spike while ignoring the spiking frequency could potentially cause them to learn overlapping representations of multiple input patterns sharing common features. We present STDP-based enhanced plasticity mechanisms that account for the spiking frequency to achieve efficient synaptic learning. First, we utilize low-pass filtered neuronal membrane potential to obtain an estimate of the spiking frequency. We perform STDP-driven weight updates in the event of a post-spike if the filtered potential exceeds a definite threshold. This ensures that plasticity is effected on the dominantly firing neuron that indicates a strong bias in learning the input pattern. Synaptic updates are restrained in the case of sporadic neuronal spiking activity, which implies a weak correlation with the input pattern. This enhances the quality of features encoded by the synapses, resulting in an improvement of 5.8% in the classification accuracy of an SNN of 100 neurons trained for digit recognition. Our simulations further show that the enhanced scheme provides a reduction of 2 χ in the number of weight updates, which leads to improved energy efficiency in event-driven SNN implementations. Second, we explore a neuronal spike-count based enhanced plasticity mechanism. The synapses are modified at the instant of a post-spike if the neuron had fired a certain number of spikes since the preceding update instant. This scheme performs delayed updates at suitable neuronal spiking instants to learn improved synaptic representations. Using this technique, the classification accuracy increased by 4% with 5.2× reduction in the number of weight updates.
Gopalakrishnan Srinivasan, Sourjya Roy, Vijay Raghunathan, Kaushik Roy 0001
IJCNN1
2016 Invited - Cross-layer approximations for neuromorphic computing: from devices to circuits and systems
abstract
Neuromorphic algorithms are being increasingly deployed across the entire computing spectrum from data centers to mobile and wearable devices to solve problems involving recognition, analytics, search and inference. For example, large-scale artificial neural networks (popularly called deep learning) now represent the state-of-the art in a wide and ever-increasing range of video/image/audio/text recognition problems. However, the growth in data sets and network complexities have led to deep learning becoming one of the most challenging workloads across the computing spectrum. We posit that approximate computing can play a key role in the quest for energy-efficient neuromorphic systems. We show how the principles of approximate computing can be applied to the design of neuromorphic systems at various layers of the computing stack. At the algorithm level, we present techniques to significantly scale down the computational requirements of a neural network with minimal impact on its accuracy. At the circuit level, we show how approximate logic and memory can be used to implement neurons and synapses in an energy-efficient manner, while still meeting accuracy requirements. A fundamental limitation to the efficiency of neuromorphic computing in traditional implementations (software and custom hardware alike) is the mismatch between neuromorphic algorithms and the underlying computing models such as von Neumann architecture and Boolean logic. To overcome this limitation, we describe how emerging spintronic devices can offer highly efficient, approximate realization of the building blocks of neuromorphic computing systems.
Priyadarshini Panda, Abhronil Sengupta, Syed Shakib Sarwar, Gopalakrishnan Srinivasan, Swagath Venkataramani, Anand Raghunathan, Kaushik Roy 0001
DAC4
2016 Significance driven hybrid 8T-6T SRAM for energy-efficient synaptic storage in artificial neural networks
Gopalakrishnan Srinivasan, Parami Wijesinghe, Syed Shakib Sarwar, Akhilesh Jaiswal 0001, Kaushik Roy 0001
DATE1