Abhiroop Bhattacharjee

dblp:273/4320 · DBLP profile ↗
← Back
23ranked-venue papers
13as first author
23since 2021 · last 2026
0000-0002-7721-271XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 12 first-author · 21 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Ultrafast Generative AI by Ultradense 3D Integration: A Case Study on LLM-based Edge Inference
abstract
Generative AI (GenAI) is one of the most critical applications today, continually challenging the limits of semiconductor technology. We introduce a very fine-grained 3D memory-on-logic architecture along with a novel data mapping strategy to support Large Language Model (LLM)-based GenAI, including both prefill and generation stages. Our conceptual analysis shows how ultradense 3D connectivity can enhance text generation speed and energy-efficiency well-beyond current limits. Preliminary findings from a basic analytical model indicate that the single batch autoregressive generation rate for Llama 3.2 1B could surpass 5K tokens/sec by maximizing weight locality and enhancing memory bandwidth through massively parallel 3D links between Multiply-Accumulate (MAC) units in the logic tier and their dedicated memory partitions in the 3D stack. We also explore the impact of advanced logic nodes and quantify their benefits in reducing prefill latency. Finally, we examine the challenges associated with memory access power and power density under extreme bandwidth conditions and present pipelined access strategies to address them.
Kerem Akarvardar, Xiaoyu Sun 0001, Brian Crafton, Xiaochen Peng, Haruki Mori, Abhiroop Bhattacharjee, Hidehiro Fujiwara, H.-S. Philip Wong
ACM Trans. Design Autom. Electr. Syst.6
2025 Low-power Spike-based Wearable Analytics on RRAM Crossbars
abstract
This work introduces a spike-based wearable analytics system utilizing Spiking Neural Networks (SNNs) deployed on an In-memory Computing engine based on RRAM crossbars, which are known for their compactness and energy-efficiency. Given the hardware constraints and noise characteristics of the underlying RRAM crossbars, we propose online adaptation of pre-trained SNNs in real-time using Direct Feedback Alignment (DFA) against traditional backpropagation (BP). Direct Feedback Alignment (DFA) learning, that allows layer-parallel gradient computations, acts as a fast, energy & area-efficient method for online adaptation of SNNs on RRAM crossbars, unleashing better algorithmic performance against those adapted using BP. Through extensive simulations using our in-house hardware evaluation engine called DFA_Sim, we find that DFA achieves upto 64.1% lower energy consumption, 10.1% lower area overhead, and a 2.1× reduction in latency compared to BP, while delivering upto 7.55% higher inference accuracy on human activity recognition (HAR) tasks.
Abhiroop Bhattacharjee, Jinquan Shi, Wei-Chen Chen, Priyadarshini Panda
ISCAS1
2025 SITRA: Exploiting Temporal Silence in Spiking Transformers for Fast & Energy-efficient Inference
abstract
Spiking Neural Network (SNN) transformers are emerging as a compelling alternative to conventional Artificial Neural Network (ANN) transformers, offering enhanced energy-efficiency through binary spike-based computations and high temporal sparsity. However, a key limitation of SNN transformers is their reliance on sequential processing across multiple timesteps, resulting in a latency drawback. This work introduces SITRA, a hardware evaluation engine tailored for SNN transformers, which employs timestep-parallelization to eliminate the latency bottleneck. Furthermore, SITRA dynamically exploits the inherent sparsity of silent inputs and preemptively detects silent outputs during SNN transformer inference. This enables skipping unnecessary costly memory accesses and operations, thereby achieving significant energy-efficiency and latency improvements over ANN transformers. Using SITRA, we find upto 6.4× higher energy-efficiency and 3.4× lower inference latency when benchmarking a SpikingBERT language transformer against its ANN baseline.
Abhiroop Bhattacharjee, Abhishek Moitra, Ruokai Yin, Priyadarshini Panda
ISLPED1
2025 ClipFormer: Key-Value Clipping of Transformers on Memristive Crossbars for Write Noise Mitigation
abstract
Transformers have revolutionized various real-world applications from natural language processing to computer vision. However, traditional von-Neumann computing paradigm faces memory and bandwidth limitations in accelerating transformers owing to their massive model sizes. To this end, in-memory computing (IMC) crossbars based on nonvolatile memories (NVMs), due to their ability to perform highly parallelized matrix-vector multiplications (MVMs) with high energy efficiencies, have emerged as a promising solution for accelerating transformers. However, analog MVM operations in crossbars introduce nonidealities, such as stochastic read and write noise, which affect the inference accuracy of the deployed transformers. Specifically, we find pretrained vision transformers (ViTs) to be vulnerable on crossbars due to the impact of write noise on the dynamically generated key$(K)$and value$(V)$matrices in the attention layers, an effect not accounted for in prior studies. We, thus, propose ClipFormer, a transformation on the K and V matrices during inference, to boost the nonideal accuracies of pretrained ViT models. ClipFormer requires no additional hardware and training overhead and is amenable to transformers deployed on any memristive crossbar platform impacted by write nonidealities. Our experiments on Imagenet-1k dataset using pretrained DeiT-S transformers, subjected to standard training and variation-aware training, show >10%–40% higher nonideal accuracies at the high write noise regime by applying ClipFormer.
Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
abstract
The attention module in vision transformers(ViTs) performs intricate spatial correlations, contributing significantly to accuracy and delay. It is thereby important to modulate the number of attentions according to the input feature complexity for optimal delay-accuracy tradeoffs. To this end, we propose PIVOT - a co-optimization framework which selectively performs attention skipping based on the input difficulty. For this,PIVOT employs a hardware-in-loop co-search to obtain optimal attention skip configurations. Evaluations on the ZCU102 MPSoC FPGA show that PIVOT achieves 2.7× lower EDP at 0.2% accuracy reduction compared to LVViT-S ViT. PIVOT also achieves 1.3% and 1.8× higher accuracy and throughput than prior works on traditional CPUs and GPUs. The PIVOT project can be found at this Github link.
Abhishek Moitra, Abhiroop Bhattacharjee, Priyadarshini Panda
DAC2
2024 HaLo-FL: Hardware-Aware Low-Precision Federated Learning
abstract
Applications of federated learning involve devices with extremely limited computational resources and often with considerable heterogeneity in terms of energy efficiency, latency tolerance, and hardware area. Although low-precision training methods have demonstrated effectiveness in accommodating the device constraints in a centralized setting, their applicability in distributed learning scenarios featuring heterogeneous client capabilities has not been well explored. In this work, we design a hardware-aware low-precision federated training framework (HaLo- FL) tailored to heterogeneous resource-constrained de-vices. In particular, we optimize the precision for weights, activations, and errors for each client's hardware constraint using a precision selector (named HaLo-PS). To validate our approach, we propose HaLoSim, a hardware evaluation platform that enables precision reconfigurability and evaluates hardware metrics like energy, latency, and area utilization on a crossbar-based In-memory Computing (IMC) platform.
Yeshwanth Venkatesha, Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda
DATE2
2024 Are SNNs Truly Energy-efficient? - A Hardware Perspective
abstract
Spiking Neural Networks (SNNs) have gained attention for their energy-efficient machine learning capabilities, utilizing bio-inspired activation functions and sparse binary spike-data representations. While recent SNN algorithmic advances achieve high accuracy on large-scale computer vision tasks, their energy-efficiency claims rely on certain impractical estimation metrics. This work studies two hardware benchmarking platforms for large-scale SNN inference, namely SATA and SpikeSim. SATA is a sparsity-aware systolic-array accelerator, while SpikeSim evaluates SNNs implemented on In-Memory Computing (IMC) based analog crossbars. Using these tools, we find that the actual energy-efficiency improvements of recent SNN algorithmic works differ significantly from their estimated values due to various hardware bottlenecks. We identify and addresses key roadblocks to efficient SNN deployment on hardware, including repeated computations & data movements over timesteps, neuronal module overhead and vulnerability of SNNs towards crossbar non-idealities.
Abhiroop Bhattacharjee, Ruokai Yin, Abhishek Moitra, Priyadarshini Panda
ICASSP1
2024 MCAIMem: A Mixed SRAM and eDRAM Cell for Area and Energy-Efficient On-Chip AI Memory
abstract
AI chips commonly employ SRAM memory as buffers for their reliability and speed, which contribute to high performance. However, SRAM is expensive and demands significant area and energy consumption. Previous studies have explored replacing SRAM with emerging technologies, such as nonvolatile memory, which offers fast read memory access and a small cell area. Despite these advantages, nonvolatile memory’s slow write memory access and high write energy consumption prevent it from surpassing SRAM performance in AI applications with extensive memory access requirements. Some research has also investigated embedded dynamic random access memory (eDRAM) as an area-efficient on-chip memory with similar access times as SRAM. Still, refresh power remains a concern, leaving the trade-off among performance, area, and power consumption unresolved. To address this issue, this article presents a novel mixed CMOS cell memory design that balances performance, area, and energy efficiency for AI memory by combining SRAM and eDRAM cells. We consider the proportion ratio of one SRAM and seven eDRAM cells in the memory to achieve area reduction using mixed CMOS cell memory. In addition, we capitalize on the characteristics of deep neural network (DNN) data representation and integrate asymmetric eDRAM cells to lower energy consumption. To validate our proposed MCAIMem solution, we conduct extensive simulations and benchmarking against traditional SRAM. Our results demonstrate that the MCAIMem significantly outperforms these alternatives in terms of area and energy efficiency. Specifically, our MCAIMem can reduce the area by 48% and energy consumption by$3.4\times $compared with SRAM designs, without incurring any accuracy loss.
Duy-Thanh Nguyen, Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda
IEEE Trans. Very Large Scale Integr. Syst.2
2023 XPert: Peripheral Circuit & Neural Architecture Co-search for Area and Energy-efficient Xbar-based Computing
abstract
The hardware-efficiency and accuracy of Deep Neural Networks (DNNs) implemented on In-memory Computing (IMC) architectures primarily depend on the DNN architecture and the peripheral circuit parameters. It is therefore essential to holistically co-search the network and peripheral parameters to achieve optimal performance. To this end, we propose XPert, which co-searches network architecture in tandem with peripheral parameters such as the type and precision of analog-to-digital converters, crossbar column sharing and the layer-specific input precision using an optimization-based design space exploration. Compared to VGG16 baselines, XPert achieves 10.24× (4.7×) lower EDAP, 1.72× (1.62×) higher TOPS/W, 1.93× (3×) higher TOPS/mm2at 92.46% (56.7%) accuracy for CIFAR10 (TinyImagenet) datasets. The code for this paper is available at the following Github link
Abhishek Moitra, Abhiroop Bhattacharjee, Youngeun Kim, Priyadarshini Panda
DAC2
2023 DeepCAM: A Fully CAM-based Inference Accelerator with Variable Hash Lengths for Energy-efficient Deep Neural Networks
abstract
With ever increasing depth and width in deep neural networks to achieve state-of-the-art performance, deep learning computation has significantly grown, and dot-products remain dominant in overall computation time. Most prior works are built on conventional dot-product where weighted input summation is used to represent the neuron operation. However, another implementation of dot-product based on the notion of angles and magnitudes in the Euclidean space has attracted limited attention. This paper proposes DeepCAM, an inference accelerator built on two critical innovations to alleviate the computation time bottleneck of convolutional neural networks. The first innovation is an approximate dot-product built on computations in the Euclidean space that can replace addition and multiplication with simple bit-wise operations. The second innovation is a dynamic size content addressable memory-based (CAM-based) accelerator to perform bit-wise operations and accelerate the CNNs with a lower computation time. Our experiments on benchmark image recognition datasets demonstrate that DeepCAM is up to$523\times$and$3498\times$faster than Eyeriss and traditional CPUs like Intel Skylake, respectively. Furthermore, the energy consumed by our DeepCAM approach is$2.16\times$to$109\times$less compared to Eyeriss.
Duy-Thanh Nguyen, Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda
DATE2
2023 Examining the Role and Limits of Batchnorm Optimization to Mitigate Diverse Hardware-noise in In-memory Computing
abstract
In-Memory Computing (IMC) platforms such as analog crossbars are gaining focus as they facilitate the acceleration of low-precision Deep Neural Networks (DNNs) with high area- & compute-efficiencies. However, the intrinsic non-idealities in crossbars, which are often non-deterministic and non-linear, degrade the performance of the deployed DNNs. In addition to quantization errors, most frequently encountered non-idealities during inference include crossbar circuit-level parasitic resistances and device-level non-idealities such as stochastic read noise and temporal drift. In this work, our goal is to closely examine the distortions caused by these non-idealities on the dot-product operations in analog crossbars and explore the feasibility of a nearly training-less solution via crossbar-aware fine-tuning of batchnorm parameters in real-time to mitigate the impact of the non-idealities. This enables reduction in hardware costs in terms of memory and training energy for IMC noise-aware retraining of the DNN weights on crossbars.
Abhiroop Bhattacharjee, Abhishek Moitra, Youngeun Kim, Yeshwanth Venkatesha, Priyadarshini Panda
ACM Great Lakes Symposium on VLSI1
2023 SpikeSim: An End-to-End Compute-in-Memory Hardware Evaluation Tool for Benchmarking Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are an active research domain toward energy-efficient machine intelligence. Compared to conventional artificial neural networks (ANNs), SNNs use temporal spike data and bio-plausible neuronal activation functions such as leaky-integrate fire/integrate fire (LIF/IF) for data processing. However, SNNs incur significant dot-product operations causing high memory and computation overhead in standard von-Neumann computing platforms. To this end, in-memory computing (IMC) architectures have been proposed to alleviate the “memory-wall bottleneck” prevalent in von-Neumann architectures. Although recent works have proposed IMC-based SNN hardware accelerators, the following key implementation aspects have been overlooked: 1) the adverse effects of crossbar nonideality on SNN performance due to repeated analog dot-product operations over multiple time-steps and 2) hardware overheads of essential SNN-specific components, such as the LIF/IF and data communication modules. To this end, we propose SpikeSim, a tool that can perform realistic performance, energy, latency and area evaluation of IMC-mapped SNNs. SpikeSim consists of a practical monolithic IMC architecture called SpikeFlow for mapping SNNs. Additionally, the nonideality computation engine (NICE) and energy–latency–area (ELA) engine performs hardware-realistic evaluation of SpikeFlow-mapped SNNs. Based on 65nm CMOS implementation and experiments on CIFAR10, CIFAR100 and TinyImagenet datasets, we find that the LIF/IF neuronal module has significant area contribution$(>11\%$of the total hardware area). To this end, we propose SNN topological modifications that leads to$1.24\times $and$10\times $reduction in the neuronal module’s area and the overall energy-delay-product value, respectively. Furthermore, in this work, we perform a holistic comparison between IMC implemented ANN and SNNs and conclude that lower number of time-steps are the key to achieve higher throughput and energy-efficiency for SNNs compared to 4-bit ANNs. The code repository for the SpikeSim tool is available at Github link.
Abhishek Moitra, Abhiroop Bhattacharjee, Runcong Kuang, Yu Cao 0001, Priyadarshini Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 SATA: Sparsity-Aware Training Accelerator for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have gained huge attention as a potential energy-efficient alternative to conventional artificial neural networks (ANNs) due to their inherent high-sparsity activation. Recently, SNNs with backpropagation through time (BPTT) have achieved a higher accuracy result on image recognition tasks than other SNN training algorithms. Despite the success from the algorithm perspective, prior works neglect the evaluation of the hardware energy overheads of BPTT, due to the lack of a hardware evaluation platform for this SNN training algorithm. Moreover, although SNNs have long been seen as an energy-efficient counterpart of ANNs, a quantitative comparison between the training cost of SNNs and ANNs is missing. To address the aforementioned issues, in this work, we introduce a sparsity-aware training accelerator (SATA), a BPTT-based training accelerator for SNNs. The proposed SATA provides a simple and reconfigurable systolic-based accelerator architecture, which makes it easy to analyze the training energy for BPTT-based SNN training algorithms. By utilizing the sparsity, SATA increases its computation energy efficiency by$5.58\times $compared to the one without using sparsity. Based on SATA, we show quantitative analyses of the energy efficiency of SNN training and make a comparison between the training cost of SNNs and ANNs. The results show that, on Eyeriss-like systolic-based architecture, SNNs consume$1.27\times $more total energy with considering sparsity (spikes, gradient of firing function, and gradient of membrane potential) when compared to ANNs. We find that such high training energy cost is from time-repetitive convolution operations and data movements during backpropagation. Moreover, to propel the future SNN training algorithm design, we provide several observations on energy efficiency for different SNN-specific training parameters and propose an energy estimation framework for SNN training.
Ruokai Yin, Abhishek Moitra, Abhiroop Bhattacharjee, Youngeun Kim, Priyadarshini Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 XploreNAS: Explore Adversarially Robust and Hardware-efficient Neural Architectures for Non-ideal Xbars
abstract
Compute In-Memory platforms such as memristive crossbars are gaining focus as they facilitate acceleration of Deep Neural Networks (DNNs) with high area and compute efficiencies. However, the intrinsic non-idealities associated with the analog nature of computing in crossbars limits the performance of the deployed DNNs. Furthermore, DNNs are shown to be vulnerable to adversarial attacks leading to severe security threats in their large-scale deployment. Thus, finding adversarially robust DNN architectures for non-ideal crossbars is critical to the safe and secure deployment of DNNs on the edge. This work proposes a two-phase algorithm-hardware co-optimization approach called XploreNAS that searches for hardware efficient and adversarially robust neural architectures for non-ideal crossbar platforms. We use the one-shot Neural Architecture Search approach to train a large Supernet with crossbar-awareness and sample adversarially robust Subnets therefrom, maintaining competitive hardware efficiency. Our experiments on crossbars with benchmark datasets (SVHN, CIFAR10, CIFAR100) show up to ~8–16% improvement in the adversarial robustness of the searched Subnets against a baseline ResNet-18 model subjected to crossbar-aware adversarial training. We benchmark our robust Subnets for Energy-Delay-Area-Products (EDAPs) using the Neurosim tool and find that with additional hardware efficiency–driven optimizations, the Subnets attain ~1.5–1.6× lower EDAPs than ResNet-18 baseline.
Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda
ACM Trans. Embed. Comput. Syst.1
2023 SwitchX: Gmin-Gmax Switching for Energy-efficient and Robust Implementation of Binarized Neural Networks on ReRAM Xbars
abstract
Memristive crossbars can efficiently implementBinarized Neural Networks(BNNs) wherein the weights are stored in high-resistance states (HRS) and low-resistance states (LRS) of the synapses. We proposeSwitchXmapping of BNN weights onto ReRAM crossbars such that the impact of crossbar non-idealities, that lead to degradation in computational accuracy, are minimized. Essentially,SwitchXmaps the binary weights in such a manner that a crossbar instance comprises of more HRS than LRS synapses. We find BNNs mapped onto crossbars withSwitchXto exhibit better robustness against adversarial attacks than the standard crossbar-mapped BNNs, the baseline. Finally, we combineSwitchXwith state-aware training (that further increases the feasibility of HRS states during weight mapping) to boost the robustness of a BNN on hardware. We find that this approach yields stronger defense against adversarial attacks than adversarial training, a state-of-the-art software defense. We perform experiments on a VGG16 BNN with benchmark datasets (CIFAR-10, CIFAR-100 and TinyImagenet) and useFast Gradient Sign Method(ε = 0.05 to 0.3) andProjected Gradient Descent( \(\epsilon =\frac{2}{255}\) to \(\frac{32}{255},~\alpha =\frac{2}{255}\) ) adversarial attacks. We show thatSwitchXcombined with state-aware training can yield upto ∼35% improvements in clean accuracy and ∼6–16% in adversarial accuracies against conventional BNNs. Furthermore, an important by-product ofSwitchXmapping is increased crossbar power savings, owing to an increased proportion of HRS synapses, which is furthered with state-aware training. We obtain upto ∼21–22% savings in crossbar power consumption for state-aware trained BNN mapped viaSwitchXon 16 × 16 and 32 × 32 crossbars using the CIFAR-10 and CIFAR-100 datasets.
Abhiroop Bhattacharjee, Priyadarshini Panda
ACM Trans. Design Autom. Electr. Syst.1
2022 MIME: adapting a single neural network for multi-task inference with memory-efficient dynamic pruning
abstract
Recent years have seen a paradigm shift towards multi-task learning. This calls for memory and energy-efficient solutions for inference in a multi-task scenario. We propose an algorithm-hardware co-design approach called MIME. MIME reuses the weight parameters of a trained parent task and learns task-specific threshold parameters for inference on multiple child tasks. We find that MIME results in highly memory-efficient DRAM storage of neural-network parameters for multiple tasks compared to conventional multi-task inference. In addition, MIME results in input-dependent dynamic neuronal pruning, thereby enabling energy-efficient inference with higher throughput on a systolic-array hardware. Our experiments with benchmark datasets (child tasks)- CIFAR10, CIFAR100, and Fashion-MNIST, show that MIME achieves ~ 3.48x memory-efficiency and ~ 2.4 - 3.1x energy-savings compared to conventional multi-task inference in Pipelined task mode.
Abhiroop Bhattacharjee, Yeshwanth Venkatesha, Abhishek Moitra, Priyadarshini Panda
DAC1
2022 Examining and Mitigating the Impact of Crossbar Non-idealities for Accurate Implementation of Sparse Deep Neural Networks
abstract
Recently several structured pruning techniques have been introduced for energy-efficient implementation of Deep Neural Networks (DNNs) with lesser number of crossbars. Although, these techniques have claimed to preserve the accuracy of the sparse DNNs on crossbars, none have studied the impact of the inexorable crossbar non-idealities on the actual performance of the pruned networks. To this end, we perform a comprehensive study to show how highly sparse DNNs, that result in significant crossbar-compression-rate, can lead to severe accuracy losses compared to unpruned DNNs mapped onto non-ideal crossbars. We perform experiments with multiple structured-pruning approaches (such as, C/F pruning, XCS and XRS) on VGG11 and VGG16 DNNs with benchmark datasets (CIFAR10 and CIFAR100). We propose two mitigation approaches - Crossbar-column rearrangement and Weight-Constrained-Training (WCT) - that can be integrated with the crossbar-mapping of the sparse DNNs to minimize accuracy losses incurred by the pruned models. These help in mitigating non-idealities by increasing the proportion of low conductance synapses on crossbars, thereby improving their computational accuracies.
Abhiroop Bhattacharjee, Lakshya Bhatnagar, Priyadarshini Panda
DATE1
2022 Rate Coding Or Direct Coding: Which One Is Better For Accurate, Robust, And Energy-Efficient Spiking Neural Networks?
abstract
Recent Spiking Neural Networks (SNNs) works focus on an image classification task, therefore various coding techniques have been proposed to convert an image into temporal binary spikes. Among them, rate coding and direct coding are regarded as prospective candidates for building a practical SNN system as they show state-of-the-art performance on large-scale datasets. Despite their usage, there is little attention to comparing these two coding schemes in a fair manner. In this paper, we conduct a comprehensive analysis of the two codings from three perspectives: accuracy, adversarial robustness, and energy-efficiency. First, we compare the performance of two coding techniques with various architectures and datasets. Then, we measure the robustness of the coding techniques on two adversarial attack methods. Finally, we compare the energy-efficiency of two coding schemes on a digital hardware platform. Our results show that direct coding can achieve better accuracy especially for a small number of timesteps. In contrast, rate coding shows better robustness to adversarial attacks owing to the non-differentiable spike generation process. Rate coding also yields higher energy-efficiency than direct coding which requires multi-bit precision for the first layer. Our study explores the characteristics of two codings, which is an important design consideration for building SNNs1.
Youngeun Kim, Hyoungseob Park, Abhishek Moitra, Abhiroop Bhattacharjee, Yeshwanth Venkatesha, Priyadarshini Panda
ICASSP4
2022 Examining the Robustness of Spiking Neural Networks on Non-ideal Memristive Crossbars
abstract
Spiking Neural Networks (SNNs) have recently emerged as the low-power alternative to Artificial Neural Networks (ANNs) owing to their asynchronous, sparse, and binary information processing. To improve the energy-efficiency and throughput, SNNs can be implemented on memristive crossbars where Multiply-and-Accumulate (MAC) operations are realized in the analog domain using emerging Non-Volatile-Memory (NVM) devices. Despite the compatibility of SNNs with memristive crossbars, there is little attention to study on the effect of intrinsic crossbar non-idealities and stochasticity on the performance of SNNs. In this paper, we conduct a comprehensive analysis of the robustness of SNNs on non-ideal crossbars. We examine SNNs trained via learning algorithms such as, surrogate gradient and ANN-SNN conversion. Our results show that repetitive crossbar computations across multiple time-steps induce error accumulation, resulting in a huge performance drop during SNN inference. We further show that SNNs trained with a smaller number of time-steps achieve better accuracy when deployed on memristive crossbars.
Abhiroop Bhattacharjee, Youngeun Kim, Abhishek Moitra, Priyadarshini Panda
ISLPED1
2022 NEAT: Nonlinearity Aware Training for Accurate, Energy-Efficient, and Robust Implementation of Neural Networks on 1T-1R Crossbars
abstract
In this era of IoT, energy-efficient and adversarially secure implementation of deep neural networks (DNNs) on hardware has become imperative. Memristive crossbars have emerged as an energy-efficient component of deep learning hardware accelerators due to their compact and efficient matrix-vector multiplication (MVM) implementation. However, they suffer from nonidealities (such as, interconnect parasitics, device variations, and sneak paths) introduced by their circuit topology that degrades computational accuracy. A 1T-1R synapse, adding a transistor (1T) in series with the memristive synapse (1R), has been proposed to mitigate sneak paths in a crossbar. However, we observe that the nonlinear characteristics of the transistor affect the overall conductance of the 1T-1R cell which in turn affects the MVM operation. This 1T-1R nonlinearity arising from the input voltage-dependent nonlinearity is not only difficult to model or formulate, but also causes a drastic performance degradation of DNNs when mapped to such crossbars. In this article, we first analyses the nonlinearity in ideal 1T-1R crossbars (excluding nonidealities, such as device variations and interconnect parasitics) and propose a novel nonlinearity aware training (NEAT) method to address the nonlinearities. Specifically, we first identify the range of network weights, which can be mapped into the 1T-1R cell within the linear operating region of the transistor. After that, we regularize the weights of neural networks to exist within the linear operating range by using an iterative training algorithm. Our iterative training significantly recovers the classification accuracy drop caused by the nonlinearity. Moreover, we find that each layer has a different weight distribution and in turn requires different gate voltage of transistor to guarantee linear operation. Based on this observation, we achieve energy efficiency while preserving classification accuracy by applying heterogeneous gate-voltage control to the 1T-1R cells across different layers. Finally, we conduct various experiments on CIFAR10 and CIFAR100 benchmark datasets to demonstrate the effectiveness of our NEAT. Overall, NEAT yields$\sim 20\%$energy gain with less than 1% accuracy loss (with homogeneous gate control) when mapping ResNet18 networks on 1T-1R crossbars. Thereafter, we integrate the 1T-1R crossbars with various nonidealities. We show that NEAT leads to more adversarially robust mappings of DNNs onto nonideal 1T-1R crossbars than standard DNNs mapped directly onto 1R crossbars. In case of a VGG11 network on CIFAR100 dataset, we obtain$\sim 17$% improvement in clean accuracy and$\sim 2$%–8% &$\sim 5$%–6% improvements in adversarial accuracies, respectively, for fast gradient sign method (FGSM) and projected gradient descent (PGD)-based adversarial attacks via NEAT on nonideal 64$\times $64 crossbars, in comparison to standard DNNs.
Abhiroop Bhattacharjee, Lakshya Bhatnagar, Youngeun Kim, Priyadarshini Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Efficiency-driven Hardware Optimization for Adversarially Robust Neural Networks
abstract
With a growing need to enable intelligence in embedded devices in the Internet of Things (IoT) era, secure hardware implementation of Deep Neural Networks (DNNs) has become imperative. We will focus on how to address adversarial robustness for DNNs through efficiency-driven hardware optimizations. Since memory (specifically, dot-product operations) is a key energy-spending component for DNNs, hardware approaches in the past have focused on optimizing the memory. One such approach is approximate digital CMOS memories with hybrid 6T-8T SRAM cells that enable supply voltage (Vdd) scaling yielding low-power operation, without significantly affecting the performance due to read/write failures incurred in the 6T cells. In this paper, we show how the bit-errors in the 6T cells of hybrid 6T-8T memories minimize the adversarial perturbations in a DNN. Essentially, we find that for different configurations of 8T-6T ratios and scaledVdd operation, noise incurred in the hybrid memory architectures is bound within specific limits. This hardware noise can potentially interfere in the creation of adversarial attacks in DNNs yielding robustness. Another memory optimization approach involves using analog memristive crossbars that perform Matrix-Vector-Multiplications (MVMs) efficiently with low energy and area requirements. However, crossbars generally suffer from intrinsic non-idealities that cause errors in performing MVMs, leading to degradation in the accuracy of the DNNs. We will show how the intrinsic hardware variations manifested through crossbar non-idealities yield adversarial robustness to the mapped DNNs without any additional optimization.
Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda
DATE1
2021 Activation Density based Mixed-Precision Quantization for Energy Efficient Neural Networks
abstract
As neural networks gain widespread adoption in embedded devices, there is a growing need for model compression techniques to facilitate seamless deployment in resource-constrained environments. Quantization is one of the go-to methods yielding state-of-the-art model compression. Most quantization approaches take a fully trained model, then apply different heuristics to determine the optimal bit-precision for different layers of the network, and finally retrain the network to regain any drop in accuracy. Based on Activation Density-the proportion of non-zero activations in a layer-we propose a novel in-training quantization method. Our method calculates optimal bit-width/precision for each layer during training yielding an energy-efficient mixed precision model with competitive accuracy. Since we train lower precision models progressively during training, our approach yields the final quantized model at lower training complexity and also eliminates the need for re-training. We run experiments on benchmark datasets like CIFAR-10, CIFAR-100, TinyImagenet on VGG19/ResNet18 architectures and report the accuracy and energy estimates for the same. We achieve up to 4.5× benefit in terms of estimated multiply-and-accumulate (MAC) reduction while reducing the training complexity by 50% in our experiments. To further evaluate the energy benefits of our proposed method, we develop a mixed-precision scalable Process In Memory (PIM) hardware accelerator platform. The hardware platform incorporates shift-add functionality for handling multibit precision neural network models. Evaluating the quantized models obtained with our proposed method on the PIM platform yields about 5× energy reduction compared to baseline 16-bit models. Additionally, we find that integrating activation density based quantization with activation density based pruning (both conducted during training) yields up to ~ 198× and ~44× energy reductions for VGG19 and ResNet18 architectures respectively on PIM platform compared to baseline 16-bit precision, unpruned models.
Karina Vasquez, Yeshwanth Venkatesha, Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda
DATE3
2021 Metastability with Emerging Reconfigurable Transistors: Exploiting Ambipolarity for Throughput
abstract
In this work, we leverage ambipolar transistors in the context of metastability for random number generation. We propose designs of a Minority-based SR latch and a dual-edge triggered True Single Phase Clock D-Flip-Flop (TSPC DFF) to sample two random bits in a single clock cycle. We demonstrate how metastable circuits based on ambipolar transistors allow doubling the throughput as compared to a similar standard CMOS-based design. The proposed design is compact in terms of the number of transistors per block (60% less transistors), power consumption (saving 94.5% leakage power and 70.7% dynamic power) and path delay (77.3% reduction) with respect to its CMOS counterpart.
Abhiroop Bhattacharjee, Shubham Rai, Ansh Rupani, Michael Raitza, Akash Kumar 0001
VLSI-SoC1