Akul Malhotra

dblp:250/9111 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-6152-2377ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SCION: A Comprehensive Simulation Framework for Charge-based In-Memory Computing for Rapid Evaluation of Hardware Non-Idealities and DNN
abstract
Charge-based in-memory computing (IMC) has shown great potential in achieving higher computational robustness compared to current-based IMC. However, it suffers from its own non-idealities such as parasitic capacitive coupling. Accurately evaluating these effects requires time-intensive SPICE simulations, making it challenging to conduct cross-layer exploration. To overcome these limitations, we propose SCION, a PyTorch-based framework that rigorously models hardware non-idealities in charge-based IMC and integrates them directly with DNN inference for rapid accuracy evaluation. We show that SCION predicts the IMC output with more than 99% accuracy with respect to SPICE while offering four orders of magnitude speedup. We demonstrate the capability of our framework by analyzing an SRAM-based charge-IMC accelerator deploying ResNet-50 and ViT-small DNNs with CIFAR-100 dataset. We show how inter-column capacitive coupling leads to data-dependent non-idealities, severely impairing the inference accuracy. We also explore techniques to mitigate non-idealities using SCION. To that end, we propose alternate column enablement (ACE) to eliminate inter-column coupling. Our results show that compared to the baseline design and another non-ideality mitigation approach based on ground shielding, ACE achieves significant improvement in sense margin and near-software accuracy under nominal conditions. Further, compared to ground shielding, ACE exhibits a higher tolerance to analog-to-digital converter (ADC) noise and superior voltage scalability.
Doug Hyun Kim, Akul Malhotra, Sumeet Kumar Gupta
ASP-DAC2
2026 X-CODA: Xbar-Output-Dependent Compensation of In-Memory-Computing Errors for DNN Accuracy Enhancement
Adrija Chakraborty, Akul Malhotra, Sumeet Kumar Gupta
ISLPED2
2025 CREST-CiM: Cross-Coupling-Enhanced Differential STT-MRAM for Robust Computing-in-Memory in Binary Neural Networks
abstract
We propose CREST-CiM, an STT-MRAM-based Computing-in-Memory (CiM) technique, targeted for binary neural networks. To circumvent the low-distinguishability issue in standard MRAM-based CiM, CREST-CiM utilizes two magnetic tunnel junctions (MTJs) to store +1 and -1 weights in a bitcell and cross-couples the MTJs, achieving a high-to-low current ratio of up to 8100 for a bit-cell. Our analysis for $64 \times 64$ arrays shows up to 3.4 x higher CiM sense-margin, 27.6% higher read-disturb-margin, and resilience to process variations, and other hardware non-idealities, albeit at the cost of just 7.9% overall-area overhead, and $\lt1 \%$ energy and latency overhead compared to a 2T-2MTJ-CiM design. Our system-level analysis for ResNet-18 trained on CIFAR-10 shows near-sofware inference accuracy with CREST-CiM, with 10.7% improvement over 2T2MTJ baseline.
Akul Malhotra, Sumeet Kumar Gupta
DAC2
2025 TWINN: Training-Free Weight-Input Flipping for Mitigating Crossbar Non-Idealities in Binary Neural Network Accelerators
abstract
Compute-in-memory (CiM)-based binary neural network (CiM-BNN) accelerators marry the benefits of CiM and ultra-low precision quantization, making them highly suitable for edge computing. However, CiM-enabled crossbar (Xbar) arrays are plagued with hardware non-idealities like parasitic resistances and device non-linearities that impair inference accuracy, especially in scaled technologies. In this work, we first analyze the impact of Xbar non-idealities on the inference accuracy of various CiM-BNNs, establishing that the unique properties of CiM-BNNs make them more prone to hardware non-idealities compared to higher precision deep neural networks (DNNs). To address this issue, we propose TWINN, a training-free technique that mitigates non-idealities in CiM-BNNs. TWINN utilizes the distinct attributes of BNNs to reduce the average current generated during the CiM operations in Xbar arrays. This is achieved by statically and dynamically flipping the BNN weights and activations, respectively. This minimizes the IR drops across the parasitic resistances, drastically mitigating their impact on inference accuracy. To evaluate our technique, we conduct experiments on ResNet-18 and VGG-small CiM-BNNs designed at the 7nm technology node using 8T-SRAM and 1T-1ReRAM. Our results show that TWINN is highly effective in alleviating the impact of non-idealities, recouping the inference accuracy to near-ideal (software) levels in some cases and providing accuracy boost of up to 77.25%. These benefits are accompanied by energy reduction, albeit at the cost of mild latency/area increase.
Akul Malhotra, Sumeet Kumar Gupta
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 ReTern: Exploiting Natural Redundancy and Sign Transformations for Enhanced Fault Tolerance in Compute-in-Memory-Based Ternary LLMs
abstract
Ternary large language models (LLMs), which use ternary precision weights and 8-bit activations, have demonstrated competitive performance while significantly reducing the high computational and memory requirements of full-precision LLMs. The energy efficiency and performance of ternary LLMs can be further improved by deploying them on ternary computing-in-memory (TCiM) accelerators, thereby alleviating the von-Neumann bottleneck. However, TCiM accelerators are prone to memory stuck-at faults (SAFs) leading to degradation in model accuracy. This is particularly severe for LLMs due to their low weight sparsity. To boost SAF tolerance of TCiM accelerators, we propose ReTern that is based on 1) fault-aware sign transformations (FASTs) and 2) TCiM bitcell reprogramming exploiting their natural redundancy. The key idea is to use FAST to minimize computation errors due to SAFs in +1/−1 weights, while the natural bitcell redundancy is exploited to target SAFs in 0 weights (zero-fix). Our experiments on BitNet b1.58 700M and 3B ternary LLMs show that our technique furnishes significant fault tolerance, notably ~35% reduction in perplexity on the Wikitext dataset in the presence of faults. These benefits come at the cost of <3%, <7%, and <1% energy, latency, and area overheads, respectively.
Akul Malhotra, Sumeet Kumar Gupta
IEEE Trans. Very Large Scale Integr. Syst.1
2024 BNN-Flip: Enhancing the Fault Tolerance and Security of Compute-in-Memory Enabled Binary Neural Network Accelerators
abstract
Compute-in-memory based binary neural networks or CiM-BNNs offer high energy/area efficiency for the design of edge deep neural network (DNN) accelerators, with only a mild accuracy reduction. However, for successful deployment, the design of CiM-BNNs must consider challenges such as memory faults and data security that plague existing DNN accelerators. In this work, we aim to mitigate both these problems simultaneously by proposing BNN-Flip, a training-free weight transformation algorithm that not only enhances the fault tolerance of CiM-BNNs but also protects them from weight theft attacks. BNN-Flip inverts the rows and columns of the BNN weight matrix in a way that reduces the impact of memory faults on the CiM-BNN’s inference accuracy, while preserving the correctness of the CiM operation. Concurrently, our technique encodes the CiM-BNN weights, securing them from weight theft. Our experiments on various CiM-BNNs show that BNN-Flip achieves an inference accuracy increase of up to 10.55% over the baseline (i.e. CiM-BNNs not employing BNN-Flip) in the presence of memory faults. Additionally, we show that the encoded weights generated by BNN-Flip furnish extremely low (near ‘random guess’) inference accuracy for the adversary attempting weight theft. The benefits of BNN-Flip come with an energy overhead of < 3%.
Akul Malhotra, Sumeet Kumar Gupta
ASPDAC1
2024 Fault Tolerant In-Memory Computing based on Emerging Technologies for Ultra-Low Precision Edge AI Accelerators
abstract
Edge Artificial Intelligence (AI) demands ultra-low power data processing on highly resource-constrained platforms, mandating a departure from conventional computing architectures. In this context, in-memory computing (IMC) coupled with ultra-low precision (ULP) neural architectures have gained traction. Furthermore, non-volatile memories such as Resistive RAMs (ReRAMs), ferroelectric-transistors (FeFETs) and others have shown an immense promise to further enhance the efficiency of deep neural network (DNN) accelerators by enabling compact and low-leakage solutions. However, the design of ULP IMC platforms must counter the impairment of inference accuracy due to manufacturing defects such as stuck-at faults (SAFs), especially those based on relatively immature emerging technologies. To that end, we present two training-free and IMC-compatible fault tolerant techniques that utilize the unique properties of ULP (binary and ternary) DNNs to mitigate the impact of SAFs. For binary neural networks (BNNs), we present BNN-Flip, a weight transformation technique that inverts rows and columns of the weight matrices to convert harmful unmasked faults to innocuous masked faults, while preserving the correctness of matrix-vector multiplication operation. Our experiments show that BNN-Flip recovers the inference accuracy of binary-precision edge devices by up to 10.55% with an energy overhead of < 3%. For ternary neural networks (TNNs), we propose TFix, a technique that exploits the natural redundancy of ternary memory arrays and high weight sparsity of TNNs to enhance fault tolerance. Our experiments show that TFix reprograms a majority of the faulty weights to their correct values, regaining the inference accuracy by up to 11.07%, with an energy overhead of < 6%.
Akul Malhotra, Sumeet Kumar Gupta
ICCAD1
2023 TFix: Exploiting the Natural Redundancy of Ternary Neural Networks for Fault Tolerant In-Memory Vector Matrix Multiplication
abstract
In-memory computing (IMC) and quantization have emerged as promising techniques for edge-based deep neural network (DNN) accelerators by reducing their energy, latency and storage requirements. In pursuit of ultra-low precision, ternary precision DNNs (TDNNs) offer high efficiency without sacrificing much inference accuracy. In this work, we explore the impact of hard faults on IMC based TDNNs and propose TFix to enhance their fault tolerance. TFix exploits the natural redundancy present in most ternary IMC bitcells as well as the high weight sparsity in TDNNs to provide up to 40.68% accuracy increase over the baseline with < 6% energy overhead.
Akul Malhotra, Sumeet Kumar Gupta
DAC1
2022 RIBoNN: Designing Robust In-Memory Binary Neural Network Accelerators
abstract
RRAM crossbar-based accelerators show promise to execute compute intensive Deep Learning applications at the edge. For highly energy-constrained systems, Binary Neural Networks (BNNs) have gained momentum in recent times as the reduced precision alleviates the costs associated with storage, compute and communication. However, faults manifested in a unit bitcell of a RRAM crossbar-based accelerator may lead to drastic degradation in accuracy of the BNN, resulting in unintended system behavior. In this paper, we propose RIBoNN, a robust RRAM-based in-memory BNN accelerator, that consists of a 2T2R differential bitcell as the basic element of the crossbar. By leveraging the inherent characteristics of the proposed bitcell, RIBoNN is capable of achieving in-situ fault tolerance, thus circumventing the need to stall the deployed application for detection or diagnosis at the edge. RIBoNN, when evaluated on image-based datasets yields up to 96.57 % improvement in BNN classification accuracy, at a fault rate of 5 %; thereby demonstrating significant fault-tolerance over the state-of-the-art XNOR-RRAM BNN accelerator. Even though RiBoNN furnishes a negligible energy overhead of 2.62% over XNOR-RRAM, our proposed accelerator significantly reduces the inference latency by performing 24.4 % faster MAC operations with identical area footprint, while providing immense fault tolerance at the edge.
Shamik Kundu, Akul Malhotra, Arnab Raha, Sumeet Kumar Gupta, Kanad Basu
ITC2