VLDB 2026 Research / reviewers in the wild / expert
Shamik Kundu
dblp:228/7028
· DBLP profile ↗
27ranked-venue papers
11as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 10 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Algorithm-Hardware Co-Design of Digital Compute-in-Memory Architecture Supporting Flexible and Temporal N:M SparsityabstractStructured pruning with fixed N:M sparsity ratios in large language models (LLMs) significantly constrains model expressivity, often leading to suboptimal accuracy. While supporting multiple N:M configurations can enhance representational flexibility, such a capability typically introduces substantial hardware complexity and overhead. To overcome these limitations, we first present FLOW, a flexible, layer-wise, outlier-density-aware N:M sparsity selection framework. FLOW adaptively determines the optimal N and M values per layer within a specified range by jointly considering the magnitude and distribution of outliers, thereby improving sparsity allocation and model fidelity. Extending this idea, we introduce FLOW++, which generalizes flexible N:M sparsity to the temporal domain for reasoning LLMs. FLOW++ enables temporal pruning by dynamically adapting sparsity patterns across reasoning thought types. To support efficient deployment of models with such dynamically varying sparsity patterns, we propose FlexCiM, a flexible, low-overhead, digital compute-in-memory (DCiM) architecture. FlexCiM partitions the DCiM macro into smaller submacros, which are adaptively aggregated and disaggregated through distribution and merging mechanisms for different values of N and M. We conduct experiments across different LLM families and state-space models and conclusively demonstrate that the proposed algorithm-hardware co-design framework achieves up to 36% higher accuracy,$1.75\times $faster inference, and$1.5\times $lower energy consumption compared to existing alternatives, establishing an effective balance between flexibility and hardware efficiency in sparse LLM inference. Code is available at:https://github.com/FLOW-open-project/FLOW Akshat Ramachandran, Souvik Kundu 0009, Arnab Raha, Shamik Kundu, Deepak Mathaikutty, Tushar Krishna |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Machine Learning-Driven STL Generation for Enhancing Functional Safety of E/E SystemsabstractThe increasing complexity of safety-critical hardware systems demands advanced methods for ensuring functional safety (FuSa). Traditional techniques like ATPG and BIST are intrusive, requiring additional hardware and disrupting operations, making them unsuitable for in-field testing. To address this, for the first time, we propose a machine learning (ML)-driven automated Self-Test Library (STL) generation for seamless in-field testing during idle periods, ensuring uninterrupted fault detection and high system performance. Utilizing reinforcement learning, the STL generates design-specific test patterns, achieving up to $57.57 \%$ improvement in fault coverage and up to $85 \%$ efficiency compared to existing pattern-based testing, enhancing FuSa in mission-critical applications. Sanjay Das, Swastik Bhattacharya, Anand Menon, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
DAC | 4 |
| 2025 | Enhancing AMS Circuit Reliability: An Anomaly Dataset for Functional Safety Research in Automotive SoCs
Sanjay Das, Anand Menon, Omar Abiola Abioye, Afreen Fatimah Khazi-Syed, Jonathan Edward Lee, Ayush Arunachalam, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
ACM Great Lakes Symposium on VLSI | 7 |
| 2025 | GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing UnitsabstractGraph Neural Networks (GNNs) are crucial for learning and reasoning over graph-structured data, with applications in network analysis, recommendation systems, and speech analytics. Deploying them on edge devices, such as client PCs and laptops, enables real-time processing, enhances privacy, and reduces cloud dependency. For instance, GNNs can augment Retrieval-Augmented Generation (RAG) for Large Language Models (LLMs) and enable event-based vision tasks. However, irregular memory access, sparse graphs, and dynamic structures lead to high latency and energy consumption on resource-constrained devices. Modern edge processors combine CPUs, GPUs, and NPUs, where NPUs excel at data-parallel tasks but face challenges with irregular GNN computations. To address these gaps, we present GraNNite, the first hardware-aware framework tailored to optimize GNN deployment on commercial-off-the-shelf (COTS) state-of-the-art (SOTA) DNN accelerators using a systematic three-step methodology: (1) enabling GNN execution on NPUs, (2) optimizing performance, and (3) trading accuracy for further performance and energy efficiency gains. Towards that end, the first category includes techniques such as GraphSplit for workload distribution and StaGr for static graph aggregation, while GrAd and NodePad handle real-time updates for dynamic graphs. Next, performance improvement is acquired through techniques such as EffOp for control-heavy operations and GraSp for sparsity exploitation. For Graph Convolution layers, PreG, SymG, and CacheG reduce redundancy and memory transfers. The final class of techniques deals with quality vs efficiency tradeoffs – QuantGr applies INT8 quantization to lower memory usage and computation time, while GrAx1, GrAx2, and GrAx3 optimize graph attention, broadcast-add, and sample-and-aggregate (SAGE)-max aggregation for higher throughput with minimal quality loss. Experimental evaluations on Intel® Core™ Ultra Series 1 and 2 AI PCs demonstrate that GraNNite achieves speedups of 2.6× to 7.6× over default NPU mappings, with energy efficiency improvements up to 8.6× compared to CPUs and GPUs. Across various GNN models, GraNNite delivers up to 10.8× and 6.7× higher performance than CPUs and GPUs, respectively. Our code implementation is available at this link. Arghadip Das, Shamik Kundu, Arnab Raha, Soumendu Kumar Ghosh, Deepak Mathaikutty, Vijay Raghunathan |
IJCNN | 2 |
| 2025 | Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory AcceleratorabstractLarge language model (LLM) pruning with fixed N:M structured sparsity significantly limits the expressivity of the sparse model, yielding sub-optimal performance. On the contrary, support for more than one N:M pattern to provide sparse representational freedom yields a costly overhead in the hardware. To mitigate these challenges for LLMs, we first present a flexible layer-wise outlier-density-aware N:M sparsity (FLOW) selection method. FLOW enables the identification of optimal layer-wise N and M values (from a given range) by simultaneously accounting for the presence and distribution of outliers, allowing a higher degree of representational freedom. To deploy the sparse models with such N:M flexibility, we then present a flexible low overhead, digital computein-memory architecture (FlexCiM). FlexCiM enables support for diverse sparsity patterns by partitioning a digital CiM (DCiM) macro into smaller sub-macros which are adaptively aggregated and disaggregated through distribution and merging mechanisms for different values of N and M. Extensive experiments on both transformer-based and recurrence-based state space foundation models (SSMs) demonstrate FLOW to outperform existing alternatives with an accuracy improvement of up to 36%, while FlexCiM delivers up to 1.75× lower inference latency and 1.5× lower energy consumption compared to existing sparse accelerators. Code is available at: https://github.com/FLOW-open-project/FLOW Akshat Ramachandran, Souvik Kundu 0009, Arnab Raha, Shamik Kundu, Deepak K. Mathaikutty, Tushar Krishna |
ISLPED | 4 |
| 2025 | OpenAssert: Towards Secure Assertion Generation using Large Language ModelsabstractAssertions are critical components used in hardware verification, ensuring robust functionality, fortifying design security, and providing essential verification features. Traditional hardware assertion methods are not automated, complicate security audits, and require effort, causing prolonged development cycles. Recent studies have highlighted the potential of commercial Large Language Models (LLMs) to generate security-focused assertions by leveraging textual data from design specifications. However, reliance on proprietary models like GPT-4 severely jeopardizes IP privacy and data confidentiality, undermining transparency and accountability in data handling practices. In this paper, we address secure hardware assertion generation by proposing a practical approach to significantly enhance the feasibility of open-source LLMs. Our proposed method, OpenAssert, involves fine-tuning existing models to be utilized locally at the user’s end without compromising confidentiality. Additionally, we employ Retrieval Augmentation Generation to refine these models, mitigating hallucinations and security-related errors. OpenAssert demonstrates improvements, achieving up to a 44% increase in rouge-1 score, a 49% improvement in cosine similarity, and a 43.4% reduction in word error rate for security-critical designs compared to open-source models. Anand Menon, Samit Shahnawaz Miftah, Amisha Srivastava, Shamik Kundu, Shovik Kundu, Arnab Raha, Suvadeep Banerjee, Deepak Mathaikutty, Kanad Basu |
VTS | 4 |
| 2024 | Graph Learning-based Fault Criticality Analysis for Enhancing Functional Safety of E/E SystemsabstractThe increasing complexity of Electrical and Electronic (E/E) systems underscores the need for protective measures to ensure functional safety (FuSa) in high-assurance environments. This entails the identification and fortification of vulnerable nodes to enhance system reliability during mission-critical scenarios. Traditionally, the assessment of E/E system reliability has relied on fault injection (FI) techniques and simulations. However, FI faces challenges in coping with escalating design complexity, including resource demands and timing overheads. Furthermore, it falls short in identifying critical components that may lead to functional failures. To address these challenges, we propose a Machine Learning (ML)-based framework for predicting critical nodes in hardware designs. The process begins with constructing a graph from the design netlist, forming the foundation for training a Graph Convolutional Network (GCN). The GCN model utilizes graph node attributes, node labels, and edge connections to learn and predict critical nodes in the circuit. The model furnishes up to 93.7% accuracy in identifying vulnerable circuit nodes during evaluation on diverse designs such as Synchronous Dynamic Random Access Memory (SDRAM) controller, OpenRISC 1200 (OR1200) modules. Furthermore, we incorporate an explainability analysis to interpret individual node predictions. This analysis discerns the critical design factors influencing fault criticality in the design. Moreover, to the best of our knowledge, we, for the first time, perform a regression analysis to generate node criticality scores, quantifying the degrees of criticality, that can enable prioritizing resources towards critical nodes. Sanjay Das, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
DAC | 2 |
| 2024 | MENDNet: Just-in-time Fault Detection and Mitigation in AI Systems with Uncertainty Quantification and Multi-Exit NetworksabstractHardware faults in AI accelerators, particularly in accelerator memory, can alter pre-trained deep neural network parameters, leading to errors that compromise performance. To address this, just-intime (JIT) fault detection and mitigation are crucial. However, existing fault detection/mitigation approaches, either interrupt continuous execution or introduce significant latency, making them less ideal for JIT implementation. To circumvent this issue, this paper explores uncertainty quantification in deep neural networks as a means of facilitating an efficient and novel fault detection approach in AI accelerators. Furthermore, in order to mitigate the impact of such faults, we propose MENDNet, which leverages the properties of multi-exit neural networks, coupled with the proposed uncertainty quantification framework. By tuning the confidence threshold for inference in each exit and leveraging the energy-based uncertainty quantification metric, MENDNet can make accurate predictions even in the presence of faults in the accelerator. When evaluated on state-of-the-art network-dataset configurations and with multiple fault rate-fault position combinations, our proposed approach furnishes up to 80.42% improvement in accuracy over a traditional DNN implementation, thereby instilling the reliability of the AI accelerator in mission mode. Shamik Kundu, Mirazul Haque, Sanjay Das, Wei Yang 0013, Kanad Basu |
DAC | 1 |
| 2024 | SwiSS: Switchable Single-Sided Sparsity-based DNN AcceleratorsabstractDeep Neural Networks (DNNs) exhibit sparsity in both activation and weight tensors, but certain layers have higher weight sparsity, while others have higher activation sparsity. This challenges the conventional approach of fixing sparsity acceleration to either weights or activations alone. Conversely, harnessing both-sided sparsity necessitates complex design logic for identifying participating non-zero weights and activation pairs during a multiply-accumulate operation, leading to a significant impact on energy efficiency and area overhead in the edge accelerator. In this paper, we, for the first time, exploit the unbalanced sparsity in DNNs to propose the concept of dynamically Switchable Single-sided Sparsity, SwiSS, to improve energy efficiency in edge DNN accelerators. Through a novel self-adaptive dynamic sparsity selection algorithm, SwiSS can determine whether to enable one-sided weight or one-sided activation sparsity for a sparsity-enabled DNN accelerator. This capability allows SwiSS to dynamically exploit both sides of sparsity while maximizing the associated power and area benefits in the accelerator. Evaluation on state-of-the-art network-dataset configurations conducted on FlexNN [12] accelerator architecture demonstrates that SwiSS yields up to 30.76% and 8.29% improvements in power and area overheads, respectively (which translates to 1.42X and 1.08X improvement in TOPS/W and TOPS/mm2, respectively), compared to a combined two-sided sparsity scenario, with a negligible drop in sparsity acceleration. Shamik Kundu, Soumendu Kumar Ghosh, Arnab Raha, Deepak Mathaikutty |
ISLPED | 1 |
| 2024 | Analyzing and Mitigating Circuit Aging Effects in Deep Learning AcceleratorsabstractThe widespread adoption of Deep Neural Networks (DNNs) can be attributed to their remarkable performance in tackling complex real-world problems. Consequently, they have found extensive use in everyday applications as well as in high-assurance environments. Nonetheless, various challenges undermine the reliability of these DNNs in mission-critical scenarios. One such challenge is circuit aging, an inevitable consequence of prolonged usage leading to the deterioration of circuit performance. Therefore, it is of utmost importance to grasp the implications of circuit aging at the application level and to adopt proactive strategies for mitigating these effects. Towards this end, our paper examines the adverse effects of circuit aging on the performance of DNN applications and introduce a novel aging-aware training (AAT) framework to mitigate such detrimental impacts. To the best of our knowledge, this framework is the first of its kind, expressly tailored to train models while considering the impact of aging. Additionally, to extend the operational lifespan of the system, as opposed to its immediate disposal, we advocate a strategic model replacement approach based on a performance threshold, particularly when aging becomes a prominent concern. Through extensive experiments involving cutting-edge DNN models, we observe substantial performance enhancements of up to 78% when utilizing AAT, even in the presence of aging, as compared to training without AAT. The model replacement approach yields significant results as well, exhibiting up to 30% relative improvement in accuracy when subjected to the same application workload. Furthermore, this improvement is augmented with AAT, achieving an additional 20% improvement, demonstrating the efficacy of the proposed framework. Sanjay Das, Shamik Kundu, Anand Menon, Yihui Ren 0001, Shubha R. Kharel, Kanad Basu |
VTS | 2 |
| 2024 | DiagNNose: Toward Error Localization in Deep Learning Hardware-Based on VTA-TVM StackabstractLow-level hardware faults manifested in a Deep learning (DL) accelerator usher in graceless degradation of high-level classification accuracy, which can eventuate to catastrophic circumstances. This violates the crucial Functional Safety (FuSa) of the DL accelerator, maintaining which is imperative in high-assurance applications. Conventional techniques for error localization incur high-test efforts, without regards to the unique challenges posed by DL systems. In this direction, we propose DiagNNose, a two-tier machine learning-based error localization framework for on-line fault management in DL accelerators. We develop a novel diagnostic pattern selection algorithm to obtain a minimal subset of functional test patterns, that are executed in the accelerator in mission mode. By extracting and analyzing dataflow-based features from the intermediate computations of the general matrix multiply (GEMM) core, a lightweight multilayer perceptron accomplishes bit-level error localization in 8-bit, 16-bit, and 32-bit datapath units with high fidelity. We have limited ourselves to a single accelerator design, i.e., the versatile tensor accelerator (VTA) architecture to evaluate our proposed DiagNNose framework. On executing state-of-the-art deep neural networks trained on ImageNet; error localization using only 30 diagnostic functional test patterns demonstrate up to 98.4% diagnosability, thereby demonstrating an improvement of 54.63% over a random test pattern set, with as low as 4.95% overhead in the DL accelerator in mission mode. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Trouble-Shooting at GAN Point: Improving Functional Safety in Deep Learning AcceleratorsabstractThe proliferation of Deep Neural Networks (DNNs) in real-time mission critical applications has promoted the implementation of custom-built DNN inference accelerators. These accelerators require a considerable amount of on-chip memory to store millions of trained DNN parameters for executing inference at the edge. Drastic technology scaling in recent years have made these memory circuits highly vulnerable to faults due to various reasons like aging, latent defects, single event upsets, etc. Such faults are highly detrimental to the classification accuracy of the DNN accelerator, leading to the crucial Functional Safety (FuSa) violation. This can eventuate to catastrophic circumstances, when used in mission-critical applications. In order to detect such violations in mission mode, we propose to generate a set of functional test patterns by leveraging the concept of Generative Adversarial Networks (GANs), that are independent of the DNN model and the accelerator characteristics. Our experimental results demonstrate that, the generated test patterns significantly improve FuSa violation detection coverage by up to 130.28%, compared to existing techniques. To the best of our knowledge, this is the first work that generates GAN-based test patterns in order to perform FuSa violation detection in mission-critical DNN accelerators. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Computers | 1 |
| 2023 | A Novel Low-Power Compression Scheme for Systolic Array-Based Deep Learning AcceleratorsabstractThe proliferation of deep learning algorithms has catalyzed their utilization to solve a multitude of real-world problems. Algorithms such as deep neural networks (DNNs) are compute- and power-intensive, thereby accentuating the development of hardware platforms like DNN inference accelerators. However, inference execution of large DNNs in resource-constrained environments induces energy bottlenecks in these accelerators. Since large DNNs consist of hundreds of millions of trained parameters, accessing them from the accelerator memory incurs substantial energy. To address this challenge, we propose HardCompress, which, to the best of our knowledge, is the first low-power solution that uses traditional compression strategies pertaining to commercial DNN accelerators in resource-constrained IoT edge devices. The three-step approach involves hardware-based post-quantization trimming of weights, followed by their dictionary-based compression and subsequent decompression by a low-power hardware engine during inference in the accelerator. We evaluate the proposed solution on lightweight networks trained on the MNIST dataset, the compact model trained on the CIFAR-10 dataset, and large DNNs trained on the ImageNet dataset. Performance of HardCompress at different quantization levels has been analyzed. Furthermore, to quantify the effectiveness of the proposed solution, an energy framework that contrasts the DRAM energies of the original and HardCompressed models has been developed. Finally, a fault injection framework which compares the fault resilience of the original model with its HardCompressed counterpart is also proposed. Our results exhibit that HardCompress, without any performance degradation in large DNNs, furnishes a maximum compression of 99.27%, equivalent to$137\times $reduction in memory footprint and 0.07 J for 8-bit quantization in the systolic array-based DNN accelerator. Furthermore, our proposed low-power decompression engine incurs an area overhead of only 0.02%; thus, enabling HardCompress’ utilization in resource-constrained environments. Ayush Arunachalam, Shamik Kundu, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | SeqL+: Secure Scan-Obfuscation With Theoretical and Empirical ValidationabstractScan-obfuscation is a powerful methodology to protect Silicon-based intellectual property from theft. Prior work on scan-obfuscation in the context of logic-locking have unique limitations, which are addressed by our previous work, SeqL, which looks at functional output corruption to obfuscate scan-chains, but is unable to resist removal attacks on circuits with inadequate number of flip-flops without feedback. To address this issue, we propose to scramble flip-flops with feedback to increase key length without introducing further vulnerabilities. This study reveals the first formulation and complexity analysis of Boolean satisfiability (SAT)-based attack on scan-scrambling. We formulate the attack as a conjunctive normal form (CNF) using a worst-case$\mathcal {O}(n^{3})$reduction in terms of scramble-graph size$n$. In order to defeat SAT-based attack, we propose an iterative swapping-based scan-cell scrambling algorithm that has$\mathcal {O}(n)$implementation time-complexity and$\mathcal {O}(2^{\lfloor ({\alpha.n+1}/{3}) \rfloor })$SAT-decryption time-complexity in terms of a user-configurable cost constraint$\alpha ~(0 < \alpha \le 1)$. Seetal Potluri, Shamik Kundu, Akash Kumar 0001, Kanad Basu, Aydin Aysu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Detecting Functional Safety Violations in Online AI AcceleratorsabstractWith the ubiquitous deployment of Deep Neural Networks (DNNs) in low latency mission critical applications, there has been an extensive proliferation of custom-built AI inference accelerators at the edge. Drastic technology scaling in recent years has made these circuits highly vulnerable to faults due to various reasons like aging, latent defects, single event upsets, etc. Such faults are highly detrimental to the classification accuracy of the AI accelerator, leading to the critical Functional Safety (FuSa) violation, when used in mission-critical applications. In order to detect such violations in mission mode, we analyze the efficiency of a software-based self test scheme that employs functional test patterns, akin to instances in the application dataset. Such patterns are either selected from the dataset of the DNN, or generated from scratch utilizing the concept of Generative Adversarial Networks (GANs). When evaluated on state-of-the-art DNNs on multivariate exhaustive datasets, the GAN generated test patterns significantly improve FuSa violation detection coverage by up to 130.28%, compared to the selected test patterns, thereby accomplishing efficient testing of the AI accelerator, online, in mission mode. Shamik Kundu, Kanad Basu |
IOLTS | 1 |
| 2022 | Design and Logic Synthesis of a Scalable, Efficient Quantum Number Theoretic TransformabstractThe advent of quantum computing has engendered a widespread proliferation of efforts utilizing qubits for optimizing classical computational algorithms. Number Theoretic Transform (NTT) is one such popular algorithm that accelerates polynomial multiplication significantly and is consequently, the core arithmetic operation in most homomorphic encryption algorithms. Hence, fast and efficient execution of NTT is highly imperative for practical implementation of homomorphic encryption schemes in different computing paradigms. In this paper, we, for the first time, propose an efficient and scalable Quantum Number Theoretic Transform (QNTT) circuit using quantum gates. We introduce a novel exponential unit for modular exponential operation, which furnishes an algorithmic complexity of O(n). Our proposed methodology performs further optimization and logic synthesis of QNTT, that is significantly fast and facilitates efficient implementations on IBM’s quantum computers. The optimized QNTT achieves a gate-level complexity reduction from power of two to one with respect to bit length. Our methodology utilizes 44.2% fewer gates, thereby minimizing the circuit depth and a corresponding reduction in overhead and error probability, for a 4-point QNTT compared to its unoptimized counterpart. Shamik Kundu, Abraham Peedikayil Kuruvila, Supriya Margabandhu Ravichandran, Kanad Basu |
ISLPED | 2 |
| 2022 | Unsupervised Learning-based Early Anomaly Detection in AMS Circuits of Automotive SoCsabstractWith the proliferation of safety-critical applications in the automotive domain, it is imperative to guarantee the functional safety of circuits and components constituting automotive systems, e.g., the electrical and/or electronic subsystems in automotive vehicles. Analog and Mixed-Signal (AMS) circuits, prevalent in such systems, are more susceptible to faults than their digital counterparts, due to advanced manufacturing nodes, parametric perturbations, environmental stress, etc. However, their continuous signal characteristics provide an opportunity for early anomaly detection, which in turn, facilitates the deployment of safety mechanisms to prevent eventual system failure. Towards this end, we propose a novel unsupervised machine learning-based framework to perform early anomaly detection in AMS circuits. Our approach involves anomaly injection in various circuit locations and individual components to develop a training dataset encompassing a wide range of possible anomalous scenarios, feature extraction from observation signals, and clustering algorithms to facilitate anomaly detection. To this end, we propose a novel centroid selection technique for the unsupervised learning algorithms, which is tailored for detecting anomalies in AMS circuits. This approach furnishes high fidelity anomaly detection by identifying the ideal cluster centers corresponding to anomalous and non-anomalous signals. Furthermore, time series-based analysis is proposed to improve and expedite the anomaly detection performance. We evaluated our solution using a case study of two AMS circuits commonly present in automotive systems-on-chips. Our experimental results exhibit that the proposed approach furnishes up to 100% accuracy. Additionally, the time series-based technique reduces the anomaly detection latency by 5×, thereby demonstrating the efficacy of our solution. Ayush Arunachalam, Athulya Kizhakkayil, Shamik Kundu, Arnab Raha, Suvadeep Banerjee, Robert Jin, Kanad Basu |
ITC | 3 |
| 2022 | RIBoNN: Designing Robust In-Memory Binary Neural Network AcceleratorsabstractRRAM crossbar-based accelerators show promise to execute compute intensive Deep Learning applications at the edge. For highly energy-constrained systems, Binary Neural Networks (BNNs) have gained momentum in recent times as the reduced precision alleviates the costs associated with storage, compute and communication. However, faults manifested in a unit bitcell of a RRAM crossbar-based accelerator may lead to drastic degradation in accuracy of the BNN, resulting in unintended system behavior. In this paper, we propose RIBoNN, a robust RRAM-based in-memory BNN accelerator, that consists of a 2T2R differential bitcell as the basic element of the crossbar. By leveraging the inherent characteristics of the proposed bitcell, RIBoNN is capable of achieving in-situ fault tolerance, thus circumventing the need to stall the deployed application for detection or diagnosis at the edge. RIBoNN, when evaluated on image-based datasets yields up to 96.57 % improvement in BNN classification accuracy, at a fault rate of 5 %; thereby demonstrating significant fault-tolerance over the state-of-the-art XNOR-RRAM BNN accelerator. Even though RiBoNN furnishes a negligible energy overhead of 2.62% over XNOR-RRAM, our proposed accelerator significantly reduces the inference latency by performing 24.4 % faster MAC operations with identical area footprint, while providing immense fault tolerance at the edge. Shamik Kundu, Akul Malhotra, Arnab Raha, Sumeet Kumar Gupta, Kanad Basu |
ITC | 1 |
| 2022 | Special Session: Effective In-field Testing of Deep Neural Network Hardware AcceleratorsabstractOngoing research to obtain high performance Deep Neural Network (DNN) executions have led to the development of customized purpose-built deep learning inference accelerators. DNN accelerators are susceptible to faults, due to high-energy particles, process variations, temperature and structural deformities manifesting as latent defects. These faults can introduce misclassification, thereby jeopardizing the Functional Safety (FuSa) of the accelerator in mission mode, which can eventuate to disastrous consequences, including loss of human lives. In this paper, we explore the impact of such faults on the FuSa of a DNN accelerator by varying the network parameters, position and characteristics of the injected fault across multiple exhaustive datasets. Furthermore, we analyze the efficiency of a software-based self test scheme to detect FuSa violations in the accelerator in mission mode, that employs functional test patterns, akin to instances in the application dataset. The test patterns, selected from the dataset of the DNN, furnish up to 100% coverage with cardinality as low as 0.1% of the entire test dataset. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Kanad Basu |
VTS | 1 |
| 2022 | Explainable Machine Learning for Intrusion Detection via Hardware Performance CountersabstractThe exponential proliferation of Malware over the past decade has threatened system security across a plethora of Internet of Things (IoT) devices. Furthermore, the improvements in computer architectures to include speculative branching and out-of-order executions have engendered new opportunities for adversaries to carry out microarchitectural attacks in these devices. Both Malware and microarchitectural attacks are imperative threats to computing systems, as their behaviors range from stealing sensitive data to total system failure. With the cat-and-mouse game between Anti-Virus Software (AVS) and attackers, the frequent bolstering of AVS induces large computational overhead. Consequently, hardware performance counter (HPC)-based detection strategies augmented with machine learning (ML) classifiers have gained popularity as a low overhead solution in identifying these malicious threats. However, ML models are operated as black boxes, which results in decisions that are not human understandable. Clarity of the models’ results facilitates the development of more robust systems. Existing explainable frameworks are only capable of determining each feature’s impact on a prediction which does not provide meaningful interpretable outcomes for HPC-based intrusion detection. In this article, we address this issue by proposing an explainable HPC-based double regression (HPCDR) ML framework. Our proposed technique provides relevant transparency through isolation of the most malevolent transient window of an application, thereby allowing a user to efficiently locate the pernicious instructions within the program. We evaluated HPCDR on five microarchitectural attacks and two Malware. HPCDR was successfully able to identify the most malicious function manifested in each intrusive application. Abraham Peedikayil Kuruvila, Shamik Kundu, Gaurav Pandey 0004, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | RTL-ConTest: Concolic Testing on RTL for Detecting Security VulnerabilitiesabstractThis article presents RTL-ConTest, a register transfer-level (RTL) security vulnerability detection algorithm, that extracts critical process flows from a RTL design and executes RTL-level concolic testing to generate security test cases for identifying critical exploits manifested in a System on Chip (SoC). The efficiency of the proposed approach is evaluated on opensource RISC-V-based SoCs. Our technique is successful in detecting the security vulnerabilities manifested in the processor core as well as in the rest of the SoC, e.g., debug modules, peripherals, etc., thereby providing a thorough vulnerability check on the entire hardware design. As demonstrated by our experimental results, in circumstances where conventional security verification tools are limited, RTL-ConTest furnishes significantly improved efficiency in detecting SoC security vulnerabilities. Shamik Kundu, Arun K. Kanuparthi, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Secure Logic Locking with Strain-Protected Nanomagnet LogicabstractPrevention of integrated circuit counterfeiting through logic locking faces the fundamental challenge of securing an obfuscation key against both physical and algorithmic threats. Previous work has focused on strengthening the logic encryption to protect the key against algorithmic attacks, but failed to provide adequate physical security. In this work, we propose a logic locking scheme that leverages the non-volatility of the nanomagnet logic (NML) family to achieve both physical and algorithmic security. Polymorphic NML minority gates protect the obfuscation key against algorithmic attacks, while a strain-inducing shield surrounding the nanomagnets provides physical security via a self-destruction mechanism. Naimul Hassan, Alexander J. Edwards, Dhritiman Bhattacharya, Mustafa M. Shihab, Varun Venkat, Peng Zhou 0025, Xuan Hu 0002, Shamik Kundu, Abraham Peedikayil Kuruvila, Kanad Basu, Jayasimha Atulasimha, Yiorgos Makris, Joseph S. Friedman |
DAC | 8 |
| 2021 | Special Session: Reliability Analysis for AI/ML HardwareabstractArtificial intelligence (AI) and Machine Learning (ML) are becoming pervasive in today's applications, such as autonomous vehicles, healthcare, aerospace, cybersecurity, and many critical applications. Ensuring the reliability and robustness of the underlying AI/ML hardware becomes our paramount importance. In this paper, we explore and evaluate the reliability of different AI/ML hardware. The first section outlines the reliability issues in a commercial systolic array-based ML accelerator in the presence of faults engendering from device-level non-idealities in the DRAM. Next, we quantified the impact of circuit-level faults in the MSB and LSB logic cones of the Multiply and Accumulate (MAC) block of the AI accelerator on the AI/ML accuracy. Finally, we present two key reliability issues- circuit aging and endurance in emerging neuromorphic hardware platforms and present our system-level approach to mitigate them. Shamik Kundu, Kanad Basu, Mehdi Sadi, Twisha Titirsha, Shihao Song, Anup Das 0001, Ujjwal Guin |
VTS | 1 |
| 2021 | Defending Hardware-Based Malware Detectors Against Adversarial AttacksabstractIn the era of Internet of Things (IoT), Malware has been proliferating exponentially over the past decade. Traditional anti-virus software are ineffective against modern complex Malware. In order to address this challenge, researchers have proposed hardware-assisted Malware detection (HMD) using hardware performance counters (HPCs). The HPCs are used to train a set of machine learning (ML) classifiers, which in turn, are used to distinguish benign programs from Malware. Recently, adversarial attacks have been designed by introducing perturbations in the HPC traces using an adversarial sample predictor to misclassify a program for specific HPCs. These attacks are designed with the basic assumption that the attacker is aware of the HPCs being used to detect Malware. Since modern processors consist of hundreds of HPCs, restricting to only a few of them for Malware detection aids the attacker. In this article, we propose a moving target defense (MTD) for this adversarial attack by designing multiple ML classifiers trained on different sets of HPCs. The MTD randomly selects a classifier; thus, confusing the attacker about the HPCs or the number of classifiers applied. We have developed an analytical model which proves that the probability of an attacker to guess the perfect HPC-classifier combination for MTD is extremely low (in the range of $10^{-1864}$ for a system with 20 HPCs). Our experimental results prove that the proposed defense is able to improve the classification accuracy of HPC traces that have been modified through an adversarial sample generator by up to 31.5%, for a near perfect (99.4%) restoration of the original accuracy. Abraham Peedikayil Kuruvila, Shamik Kundu, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Toward Functional Safety of Systolic Array-Based Deep Learning Hardware AcceleratorsabstractHigh accuracy and ever-increasing computing power have made deep neural networks (DNNs) the algorithm of choice for various machine learning, computer vision, and image processing applications across the computing spectrum. To this end, Google developed the tensor processing unit (TPU) to accelerate the computationally intensive matrix multiplication operation of a DNN on its systolic array architecture. Faults manifested in the datapath of such a systolic array due to latent manufacturing defects or single-event effects may lead to functional safety (FuSa) violation. Although DNNs are known to resist minor perturbations with their inherent fault-tolerant characteristics, we show that the classification accuracy of the model plummets from 97.4% to 7.75% with a minimal fault rate of 0.0003% in the accelerator, implying catastrophic circumstances when deployed across mission-critical systems. Hence, to ensure FuSa of such accelerators, this article provides an extensive FuSa assessment of the accelerator exposed to faults in the datapath, by varying the network parameters, position, and characteristics of the induced error across multiple exhaustive data sets. Furthermore, we propose two novel strategies to obtain a diminutive set of functional test patterns to detect FuSa violation in a DNN accelerator. Our experimental results demonstrate that the obtained test sets can achieve an average of 92.63% (in some cases, up to 100%) fault coverage with cardinality as low as 0.1% of the entire test data set. Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | High-level Modeling of Manufacturing Faults in Deep Neural Network AcceleratorsabstractThe advent of data-driven real-time applications requires the implementation of Deep Neural Networks (DNNs) on Machine Learning accelerators. Google's Tensor Processing Unit (TPU) is one such neural network accelerator that uses systolic array-based matrix multiplication hardware for computation in its crux. Manufacturing faults at any state element of the matrix multiplication unit can cause unexpected errors in these inference networks. In this paper, we propose a formal model of permanent faults and their propagation in a TPU using the Discrete-Time Markov Chain (DTMC) formalism. The proposed model is analyzed using the probabilistic model checking technique to reason about the likelihood of faulty outputs. The obtained quantitative results show that the classification accuracy is sensitive to the type of permanent faults as well as their location, bit position and the number of layers in the neural network. The conclusions from our theoretical model have been validated using experiments on a digit recognition-based DNN. Shamik Kundu, Ahmet Soyyigit, Khaza Anuarul Hoque, Kanad Basu |
IOLTS | 1 |
| 2018 | Classification of Short-Texts Generated During Disasters: A Deep Neural Network Based ApproachabstractMicro-blogging sites provide a wealth of resources during disaster events in the form of short texts. Correct classification of these text data into various actionable classes can be of great help in shaping the means to rescue people in disaster-affected places. The process of classification of these text data poses a challenging problem because the texts are usually short and very noisy and finding good features that can distinguish these texts into different classes is time consuming, tedious and often requires a lot of domain knowledge. We propose a deep learning based model to classify tweets into different actionable classes such as resource need and availability, activities of various NGO etc. Our model requires no domain knowledge and can be used in any disaster scenario with little to no modification. Shamik Kundu, P. K. Srijith, Maunendra Sankar Desarkar |
ASONAM | 1 |