Jonti Talukdar

dblp:205/2729 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0001-7079-5281ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 10 first-author · 27 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Focus Session: Do Agentic LLMs Change the Paradigm of Hardware Test Generation?
abstract
Technology scaling and increasing System-on-Chip (SoC) complexity exacerbate reliability challenges arising from both structural defects and runtime-dependent failures, including Silent Data Corruptions (SDCs) that evade traditional error detection mechanisms. Structural testing remains essential for detecting modeled faults such as stuck-at faults; however, it is inherently limited in capturing failures that arise under dynamic operating conditions. In contrast, functional testing can expose workload-dependent failures, albeit at the cost of high testing overhead and largely unguided workload generation. This paper presents an agentic testing framework that integrates Large Language Models (LLMs) with Reinforcement Learning (RL) and Tree-structured Parzen Estimators (TPE) to guide functional workload generation and Automatic Test Pattern Generation (ATPG) settings under user-defined constraints. The proposed approach leverages feedback-driven optimization to steer test generation toward failure-prone behaviors while reducing reliance on manual expertise. Experimental evaluation on a RISC-V processor core demonstrates that the method outperforms manually generated workloads for functional testing, while experiments on six benchmark circuits show test quality comparable to expert-generated ATPG scripts for structural testing, with improved efficiency and scalability.
Farshad Firouzi, Agastya Seth, Peter Domanski, Bahareh J. Farahani, Sanmitra Banerjee, Jonti Talukdar, Krishnendu Chakrabarty
DATE7
2026 LEAD: Link Exploitability Analysis for Die-to-Die Interconnects in Heterogeneous Integration*
Arjun Hati, Eduardo Ortega, Jonti Talukdar, James F. Plusquellic, Krishnendu Chakrabarty
VTS3
2026 TRACK: Telemetry-based Representation Analysis via Centered Kernel Alignment for Silicon Lifecycle Management
Eduardo Ortega, Jonti Talukdar, Hsiao-Ping Ni, Krishnendu Chakrabarty
VTS2
2026 Innovative Practices Session: Hardware Security and Test
Jonti Talukdar, Sohrab Aftabjahani
VTS1
2026 CATCH: A Cost Analysis Tool for Co-Optimization of Chiplet-Based Heterogeneous Systems
abstract
With the increasing prevalence of chiplet systems in high-performance computing applications, the number of design options has increased dramatically. Instead of chips defaulting to a single die per package, now there are viable 3D stacking options to integrate multiple dies in a package either through vertical stacking or horizontal integration on a substrate along with a plethora of choices regarding configurations and processes. For chiplet-based designs, high-impact decisions such as those regarding the number of chiplets, the design partitions, the interconnect types, and other factors must be made early in the development process. In this work, we describe an open-source tool, CATCH, that can be used to guide these early design choices. We also present case studies showing some of the insights we can draw by using this tool. We look at case studies on optimal chip size, defect density, test cost, IO types, assembly processes, and substrates. Additionally, we include the cost breakdown for a specific case based on a prior work.
Alexander Graening, Jonti Talukdar, Saptadeep Pal, Krishnendu Chakrabarty, Puneet Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 Dynamic Interposer Obfuscation Through Distributed Scramblers in Heterogeneously Integrated 2.5D ICs
abstract
Recent breakthroughs in heterogeneous integration (HI) using 2.5D and 3D ICs have been key to advances in the semiconductor industry. However, heterogeneous integration has also led to several sources of distrust due to the use of thirdparty IP, testing, and fabrication facilities in the design and manufacturing process. Recent work on 2.5D IC security has focused on attacks that can be mounted through rogue chiplets integrated in the design. Thus, existing solutions implement interchiplet communication protocols that prevent unauthorized data modification and interruption in a 2.5D system. However, none of the existing solutions offer inherent security against IP theft. We develop a comprehensive threat model for 2.5D systems indicating that such systems remain vulnerable to IP theft. We present a method that prevents IP theft by obfuscating the connectivity of chiplets on the interposer using reconfigurable interconnection networks. We further achieve dynamic obfuscation of chiplet interconnects while demonstrating resilience against removal, data-snooping, test-based data leakage attacks. We present an approach to implement reconfigurable chiplet interconnection networks through both centralized and distributed scrambler designs, evaluating the security and implementation benefits of both the architectures. We present a comprehensive methodology to design, implement, and integrate interconnect scramblers in a large 2.5D HI system. We also evaluate the power, performance, and area overhead for both distributed and centralized scramblers.
Jonti Talukdar, Seungmin Woo, Sung Kyu Lim, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 Runtime Fault Localization in Deep Neural Network Accelerators
abstract
Systolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Although fault detection and repair techniques have been proposed to enhance the robustness of systolic arrays, fault localization remains an open problem. We propose a fault tolerance framework including run-time based fault detection and fault localization, both leveraging functional data to generate checksums on-the-fly. This approach enables error detection and localization during normal operation, avoiding the need for dedicated test patterns or additional downtime. Experimental evaluation shows that the proposed fault localization architecture incurs an area overhead less than 2% for a 256× 256 systolic array. In simulations, the proposed method achieves 100% fault detection and localization in a 256× 256 systolic array.
Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty
ACM Trans. Design Autom. Electr. Syst.2
2025 MALLS: Multi-Agent LLMs for Synthetic Hardware Vulnerability Generation and Detection
abstract
LLMs have demonstrated promising capabilities in generating RTL code from high-level functional descriptions of hardware modules. However, their effectiveness is constrained by the lack of high-quality, diverse datasets particularly for applications in IP design, verification, and security analysis. To address this limitation, we introduce MALLS, a multi-agent framework in which specialized LLM agents namely, a generator and a discriminator collaborate in an adversarial yet cooperative setting to improve the quality and correctness of RTL designs and curate a high quality synthetic hardware vulnerability dataset. In this architecture, the generator agent is responsible for producing RTL implementations from initial seed examples through in-context learning, while the discriminator agent assesses the generator's output for functional correctness and the presence of security vulnerabilities. This interaction creates a dynamic feedback loop, enabling both agents to iteratively improve through each other's responses leading to self-supervised learning. A hard bank of examples is maintained in a database which include instances that were difficult to generate or detect by either agents. By generating paired positive (correct) and negative (buggy) examples, the system learns to distinguish subtle design flaws and generalize across diverse RTL patterns while generating high quality synthetic examples of hardware vulnerabilities. Experimental results show that using the hard bank of examples produced by the adversarial multi-LLM setup improves both vulnerability generation and detection performance.
Jonti Talukdar, Agastya Seth, Sanmitra Banerjee, Farshad Firouzi, Krishnendu Chakrabarty
ICCD1
2025 OCTANE: On-Chip Telemetry-based Anomaly Notification Engine
abstract
Silicon lifecycle management (SLM) is essential for ensuring the reliability and quality of silicon products. Traditional approaches primarily rely on off-chip solutions to detect malware, diagnose hardware bugs, and characterize silicon health metrics. However, these methods do not incorporate hardware/software co-design for SLM. This work introduces On- Chip Telemetry-based Anomaly Notification Engine (OCTANE), designed to monitor chip status using performance counters and sensors. OCTANE features a compute- and memory-efficient, unsupervised anomaly detection mechanism implemented on-chip (OCTANE-edge) using fixed-point arithmetic. Furthermore, it enhances on-chip anomaly detection through unsupervised feature ranking based on telemetry feature information entropy and compression index. This unsupervised feature ranking technique is workload-independent and provides the generalizability required for SLM. The proposed solution extends to an end-to-end anomaly-informed diagnosis model that leverages OCTANE-edge compacted anomaly telemetry signatures to diagnose chip security or safety incidents (OCTANE-cloud). All telemetry data is collected via model-specific register space using open-source Linux tools and the performance counter monitor. To validate our approach, we capture chip telemetry signatures from the PAMPAR benchmark suite under anomaly-inducing events such as security attacks (e.g., Rowhammer and Spectre) and voltage droops across two Intel platforms. OCTANE demonstrates highly effective unsupervised on-chip anomaly detection with minimal area and idle power overhead (1.2% and 2.6%). OCTANE provides anomaly detection/diagnosis with accuracy surpassing 0.96/0.98.
Eduardo Ortega, Arjun Hati, Jonti Talukdar, Woohyun Paik, Rita Chattopadhyay, Krishnendu Chakrabarty
ITC3
2025 TaintLock: Hardware IP Protection Against Oracle-Guided and Oracle-Reconstruction Attacks
abstract
Scan-obfuscation schemes used with logic locking lack the ability to perform scan authentication on a per-pattern basis. These methods are of limited effectiveness in obfuscating scan data and they remain vulnerable to SAT-based scan deobfuscation attacks. In addition, prior methods designed to perform scan-data authentication are not adequate under the strongest threat models used to assess logic locking. To alleviate these problems, we propose enhancements to TaintLock, a lightweight dynamic per-pattern authentication and encryption scheme that uses taint and signature bits embedded within each test pattern to provide authenticated scan access. To prevent IP theft through Oracle-free and Oracle-guided attacks, TaintLock is paired with truly random logic locking (TRLL). TaintLock cryptographically authenticates each test pattern using the embedded taint and signature bits and passing them through a substitution-permutation (SP) network. It further uses cryptographically generated keys to dynamically encrypt scan data for unauthenticated users. TaintLock, while offering a low overhead and nonintrusive secure scan solution may remain susceptible to a new class of Oracle-reconstruction attacks that use machine learning. Additionally, assuming test pattern security is compromised, it may be potentially vulnerable to a template-based SAT attack aimed at partial key recovery. We analyze the susceptibility of TaintLock against these threats and demonstrate its resilience. We also demonstrate that TaintLock can be easily integrated with popular test architectures, such as embedded deterministic test (EDT). Finally, we also discuss the reconfigurable nature of TaintLock’s architecture to support different levels of encryption and authentication.
Jonti Talukdar, Arjun Chaudhuri, Eduardo Ortega, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Detection of Voltage Droop-Induced Timing Fault Attacks Due to Hardware Trojans
abstract
Recent breakthroughs in heterogeneous integration (HI) using 2.5-D/3-D packaging technology have led to several advances in the semiconductor industry, increasing yield while reducing overall cost and time-to-market. However, the diversification of the HI supply chain has led to several sources of distrust due to the use of black-boxed third-party intellectual property (IP), outsourced fabrication, assembly and test facilities during the design and manufacturing process. We demonstrate the susceptibility of chiplet IPs to timing failure due to voltage droop in the power distribution network (PDN) induced by the insertion of chiplet level ring-oscillator (RO)-based hardware Trojans. We present an end-to-end methodology for design, placement, and insertion of RO-based Trojans in chiplet designs followed by characterizing their contribution to the dynamic voltage droop induced within the on-chip PDN. We quantify this PDN impact on timing paths and develop a systematic method to rank the susceptibility of different data paths toward a voltage droop event. We utilize this presilicon security analysis framework to evaluate voltage droop-based attack susceptibility for a variety of IPs, including some from the CEP benchmark. We also develop a machine learning-guided time-series anomaly detection framework to detect voltage droop-based anomalies on functional workloads running on different benchmarks, demonstrating the effectiveness of an convolutional autoencoders in detecting voltage droop-induced timing anomalies.
Jonti Talukdar, Akshay Vyas, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Effective Runtime Fault Detection for DNN Accelerators
abstract
Systolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing of matrix multiplication, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Even though many algorithm-based fault tolerance (ABFT) algorithms have been proposed to detect and correct errors in matrix multiplication, these ABFT methods cannot detect many errors originating from the accelerator hardware. We propose a run-time based fault detection technique leveraging functional data to generate checksums on-the-fly, avoiding the requirement for test patterns. Experimental evaluation shows that the proposed fault detection architecture can achieve 100% test coverage while incurring an area overhead of less than 2% for a 256 × 256 systolic array.
Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty
ATS2
2024 LLM-AID: Leveraging Large Language Models for Rapid Domain-Specific Accelerator Development
abstract
The challenges posed by the Dark Silicon era, combined with the escalating computational demands of emerging applications, such as Deep Learning (DL), have strained the capabilities of traditional CPUs and GPUs, necessitating the development of Domain-Specific Accelerators (DSAs). Despite offering substantial enhancements in Power, Performance, and Area (PPA), DSAs encounter significant challenges, including the rapid evolution of applications that necessitate the frequent development of new architectures. This, coupled with the expertise-intensive nature of the design process, often leads to reduced flexibility and extended development cycles, ultimately hindering the broader adoption and efficient deployment of DSAs. To address these challenges, this paper introduces LLM-AID, an agile framework that streamlines the DSA design flow by transforming high-level abstract specifications into Hardware Description Language (HDL) code and facilitating backend Computer-Aided Design (CAD) tool operations. By synergistically combining Large Language Models (LLMs), High-Level Synthesis (HLS) tools, design exploration techniques, and symbolic AI, LLM-AID dramatically accelerates design iterations, optimizes hardware performance, and significantly reduces time-to-market. This innovative approach democratizes DSA development, empowering designers to achieve unprecedented productivity while delivering high-quality DSA solutions.
Farshad Firouzi, Sri Sai Rakesh Nakkilla, Chenghao Fu, Sanmitra Banerjee, Jonti Talukdar, Krishnendu Chakrabarty
ICCAD5
2024 E-SCOUT: Efficient-Spatial Clustering-based Outlier Detection through Telemetry
abstract
Silicon lifecycle management (SLM) is needed to ensure silicon-product reliability and quality. Prior methods utilize off-chip solutions to identify malware, diagnose bugs, and characterize silicon health metrics. These methods do not explore hardware/software codesign for SLM. In this work, we present a new method called Efficient-Spatial Clustering-based OUlier detection through Telemetry (E-SCOUT) to monitor a chip’s status through performance counters/sensors. E-SCOUT includes a compute- and memory-efficient unsupervised 32-bit floating point outlier detection mechanism implemented on-chip (E-SCOUT edge). In addition, it enhances on-chip outlier detection through unsupervised feature ranking based on the telemetry feature information entropy. We also provide microarchitectural recommendations to enable a hardware/software co-design of E-SCOUT edge. The proposed solution includes an end-to-end outlier-informed diagnosis model with real telemetry data (E-SCOUT cloud). All telemetry data is collected through the model-specific register space using open-source Linux tools and Intel’s performance counter monitor. We capture the chip telemetry signatures of the PAMPAR benchmark suite in the presence of outlier events such as security attacks (e.g., Rowhammer and SPECTRE) and voltage droops. E-SCOUT provides effective unsupervised on-chip outlier detection performance with high accuracy levels (over 0.9) and with low area and low power over-head (2.2% die area overhead and 1% idle power consumption). Outlier diagnosis can identify the chip’s status with classification accuracy and F1-scores that exceed 0.8.
Eduardo Ortega, Jonti Talukdar, Woohyun Paik, Rita Chattopadhyay, Krishnendu Chakrabarty
ITC2
2024 Rowhammer Vulnerability of DRAMs in 3-D Integration
abstract
We investigate the vulnerability of 3-D-integrated dynamic random access memorys (DRAMs) [i.e., typically connected with silicon via (TSV), monolithic interconnect via (MIV)] to Rowhammer attacks. We have developed a SPICE framework to characterize Rowhammer attacks for the scenarios described. We utilize OPENROAD ASAP7 PDK for our simulation. We investigate horizontal (within the same tier) and vertical (across multiple tiers) variants of Rowhammer attacks. We show that horizontal Rowhammer vulnerability may be reduced through DRAM bank partitioning. In addition, we show that vertical parasitic capacitance in TSV 3D-DRAM is unlikely to lead to vertical Rowhammer attacks. However, vertical parasitic capacitance in MIV 3D-DRAM can make vertical Rowhammer attacks feasible.
Eduardo Ortega, Jonti Talukdar, Woohyun Paik, Tyler K. Bletsch, Krishnendu Chakrabarty
IEEE Trans. Very Large Scale Integr. Syst.2
2024 ALT-Lock: Logic and Timing Ambiguity-Based IP Obfuscation Against Reverse Engineering
abstract
We present a logic ambiguity-based intellectual property (IP) obfuscation method that replaces traditional key gates with key-controlled functionally ambiguous logic gates, called LGA gates. We also protect timing paths by developing timing-ambiguous sequential cells called TA cells. We call this locking scheme ambiguous logic and timing logic locking (referred to as ALT-Lock). ALT-Lock ensures a two-pronged system-level security scheme where the attacker is forced to unlock not only combinational logic obfuscation but also timing obfuscation. We show that a combination of logic and timing ambiguity (TA) provides security against oracle-guided attacks. This method is superior to other traditional IP protection schemes such as combinational or sequential locking as it guarantees security against both oracle-guided and oracle-free attacks, while ensuring low power, performance, and area (PPA) overhead.
Jonti Talukdar, Woohyun Paik, Eduardo Ortega, Krishnendu Chakrabarty
IEEE Trans. Very Large Scale Integr. Syst.1
2023 Securing Heterogeneous 2.5D ICs Against IP Theft through Dynamic Interposer Obfuscation
abstract
Recent breakthroughs in heterogeneous integration (HI) technologies using 2.5D and 3D ICs have been key to advances in the semiconductor industry. However, heterogeneous integration has also led to several sources of distrust due to the use of third-party IP, testing, and fabrication facilities in the design and manufacturing process. Recent work on 2.5D IC security has only focused on attacks that can be mounted through rogue chiplets integrated in the design. Thus, existing solutions implement inter-chip let communication protocols that prevent unauthorized data modification and interruption in a 2.5D system. However, none of the existing solutions offer inherent security against IP theft. We develop a comprehensive threat model for 2.5D systems indicating that such systems remain vulnerable to IP theft. We present a method that prevents IP theft by obfuscating the connectivity of chiplets on the interposer using reconfigurable interconnection networks. We also evaluate the PPA impact and security offered by our proposed scheme.
Jonti Talukdar, Arjun Chaudhuri, Sung Kyu Lim, Krishnendu Chakrabarty
DATE1
2023 Simply-Track-and-Refresh: Efficient and Scalable Rowhammer Mitigation
abstract
Rowhammer is a memory vulnerability that can compromise system-level security. Rowhammer occurs when a DRAM row is accessed repeatedly, potentially causing bit-flips for neighboring rows. The threshold for Rowhammer has decreased from 139K accesses in 2014 to 3.2K in 2022. This threshold is projected to decrease further. Many existing solutions are not scalable, incur high overhead, or fail to offer protection in realistic scenarios. We propose Simply-Track-And-Refresh (STAR) as an effective and scalable Rowhammer mitigation. We compare STAR's performance overhead to recent solutions, HYDRA and AQUA. At ultra-low thresholds (500), STAR introduces 9.5x/31.7x lower average execution time overhead than HYDRA/AQUA. In addition, STAR introduces up to 4.3x lower area overhead and up to 3.3x lower power consumption compared to HYDRA and AQUA. We present proof of correctness, area and power consumption results derived using CACTI, and evaluation results from the PARSEC, SPLASH-2, SPEC2006, SPEC2017, and PAMPAR benchmark suites.
Eduardo Ortega, Tyler K. Bletsch, Biresh Kumar Joardar, Jonti Talukdar, Woohyun Paik, Krishnendu Chakrabarty
ITC4
2023 Functional Test Generation for AI Accelerators using Bayesian Optimization∗
abstract
We propose a black-box optimization method to generate functional test patterns for AI inferencing accelerators. Functional testing is faster than structural testing as scan chains are not used for shifting in patterns and shifting out test responses. Moreover, functional testing reduces "over-testing" by targeting the detection of functionally critical faults for a given application workload. We use Bayesian Optimization for targeted test-image generation for stuck-at faults in a systolic array-based accelerator. Our framework supports test-pattern compaction and leverages various types of error regularization for enforcing functional-likeness of the generated test images. We achieve high fault coverage using a small set of test images for pin-level faults in 16-bit and 32-bit floating-point processing elements of the systolic array achieves high fault coverage with a small set of test images.
Arjun Chaudhuri, Ching-Yuan Chen, Jonti Talukdar, Krishnendu Chakrabarty
VTS3
2022 TaintLock: Preventing IP Theft through Lightweight Dynamic Scan Encryption using Taint Bits*
abstract
We propose TaintLock, a lightweight dynamic scan data authentication and encryption scheme that performs per-pattern authentication and encryption using taint and signature bits embedded within the test pattern. To prevent IP theft, we pair TaintLock with truly random logic locking (TRLL) to ensure resilience against both Oracle-guided and Oracle-free attacks, including scan deobfuscation attacks. TaintLock uses a substitution-permutation (SP) network to cryptographically authenticate each test pattern using embedded taint and signature bits. It further uses cryptographically generated keys to encrypt scan data for unauthenticated users dynamically. We show that it offers a low overhead, non-intrusive secure scan solution without impacting test coverage or test time while preventing IP theft.
Jonti Talukdar, Arjun Chaudhuri, Krishnendu Chakrabarty
ETS1
2022 Machine Learning for Testing Machine-Learning Hardware: A Virtuous Cycle
abstract
The ubiquitous application of deep neural networks (DNN) has led to a rise in demand for AI accelerators. DNN-specific functional criticality analysis identifies faults that cause measurable and significant deviations from acceptable requirements such as the inferencing accuracy. This paper examines the problem of classifying structural faults in the processing elements (PEs) of systolic-array accelerators. We first present a two-tier machine-learning (ML) based method to assess the functional criticality of faults. While supervised learning techniques can be used to accurately estimate fault criticality, it requires a considerable amount of ground truth for model training. We therefore describe a neural-twin framework for analyzing fault criticality with a negligible amount of ground-truth data. We further describe a topological and probabilistic framework to estimate the expected number of PE's primary outputs (POs) flipping in the presence of defects and use the PO-flip count as a surrogate for determining fault criticality. We demonstrate that the combination of PO-flip count and neural twin-enabled sensitivity analysis of internal nets can be used as additional features in existing ML-based criticality classifiers.
Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty
ICCAD2
2022 Automatic Structural Test Generation for Analog Circuits using Neural Twins
abstract
The growing size of analog IPs has made targeted structural testing of such designs a challenging problem. We present a gradient-based automated test generation framework for analog circuits using neural twins, which are neural equivalents of the corresponding analog circuit. A neural twin is constructed by combining several FET-twins that lie in the paths between the circuit's inputs and observation points. Each FET-twin is a fully-connected neural network that models the IV characteristics of individual MOSFETs in the design. We train different variants of FET-twins that can predict both the output current and nodal voltage with more than 99% accuracy. We create an analog neural miter circuit, for which tests are generated using gradient ascent to maximize the loss between the faulty and fault-free versions of the neural twin. By computing gradients in a batchwise fashion for all the faults in the design, we develop a test compaction scheme that covers all faults with minimum number of test patterns. The neural twin-driven test generation method is interpretable, faster to simulate through GPUs, and guarantees convergence through backpropagation. We demonstrate the effectiveness of this framework by generating tests for structural defects in analog benchmark circuits. We show that our method outperforms an existing black-box optimization method that can be repurposed for test generation.
Jonti Talukdar, Arjun Chaudhuri, Mayukh Bhattacharya, Krishnendu Chakrabarty
ITC1
2022 Special Session: Fault Criticality Assessment in AI Accelerators
abstract
The ubiquitous application of deep neural networks (DNN) has led to a rise in demand for AI accelerators. DNN-specific functional criticality analysis identifies faults that cause measurable and significant deviations from acceptable requirements such as the inferencing accuracy. This paper examines the problem of classifying structural faults in the processing elements (PEs) of systolic-array accelerators. We first present a two-tier machine-learning (ML) based method to assess the functional criticality of faults. The problem of minimizing misclassification is addressed by utilizing generative adversarial networks (GANs). The two-tier ML/GAN-based criticality assessment method leads to less than 1% test escapes during functional criticality evaluation of structural faults. While supervised learning techniques can be used to accurately estimate fault criticality, it requires a considerable amount of ground truth for model training. We therefore describe a neural-twin framework for analyzing fault criticality with a negligible amount of ground-truth data. A recently proposed misclassification-driven training algorithm is used to sensitize and identify biases that are critical to the functioning of the accelerator for a given application workload. The proposed framework achieves up to 100% accuracy in fault-criticality classification in 16-bit and 32-bit PEs by using the criticality knowledge of only 2% of the total faults in a PE.
Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty
VTS2
2022 Functional Criticality Analysis of Structural Faults in AI Accelerators
abstract
The ubiquitous application of deep neural networks (DNNs) has led to a rise in demand for artificial intelligence (AI) accelerators. For example, the tensor processing unit from Google–based on a systolic array–and its variants are of considerable interest for DNN inferencing using AI accelerators. This article studies the problem of classifying structural faults in such an accelerator based on their functional criticality. We first analyze pin-level faults in the processing elements (PEs) of a systolic array. Simulation results for the LeNet network with 8-bit fixed-point, 16-bit floating-point (FP), and 32-bit FP data paths applied to the MNIST dataset show that over 93% of the pin-level structural faults in a PE are functionally benign. We present a greedy iterative framework for determining the criticality of stuck-at faults in a PE netlist and analyze the limitations of criticality analysis methods based on repeated fault simulations. We next present a scalable two-tier machine-learning (ML)-based method to assess the functional criticality of stuck-at faults in a computationally efficient manner. We address the problem of minimizing misclassification by utilizing generative adversarial networks (GANs). Two-tier ML/GAN-based criticality assessment leads to less than 1% test escapes during functional criticality evaluation of structural faults.
Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Fault-Criticality Assessment for AI Accelerators using Graph Convolutional Networks
abstract
Owing to the inherent fault tolerance of deep neural networks (DNNs), many structural faults in DNN accelerators tend to be functionally benign. In order to identify functionally critical faults, we analyze the functional impact of stuck-at faults in the processing elements of a 128×128 systolic-array accelerator that performs inferencing on the MNIST dataset. We present a 2-tier machine-learning framework that leverages graph convolutional networks (GCNs) for quick assessment of the functional criticality of structural faults. We describe a computationally efficient methodology for data sampling and feature engineering to train the GCN-based framework. The proposed framework achieves up to 90% classification accuracy with negligible misclassification of critical faults.
Arjun Chaudhuri, Jonti Talukdar, Jinwook Jung, Gi-Joon Nam, Krishnendu Chakrabarty
DATE2
2021 Efficient Fault-Criticality Analysis for AI Accelerators using a Neural Twin∗
abstract
Owing to the inherent fault tolerance of deep neural network (DNN) models used for classification, many structural faults in the processing elements (PEs) of a systolic array-based AI accelerator are functionally benign. Brute-force fault simulation for determining fault criticality is computationally expensive due to many potential fault sites in the accelerator array and the dependence of criticality characterization of PEs on the functional input data. Supervised learning techniques can be used to accurately estimate fault criticality but it requires ground truth for model training. The ground-truth collection involves extensive and computationally expensive fault simulations. We present a framework for analyzing fault criticality with a negligible amount of ground-truth data. We incorporate the gate-level structural and functional information of the PEs in their "neural twins", referred to as "PE-Nets". The PE netlist is translated into a trainable PE-Net, where the standard-cell instances are substituted by their corresponding "Cell-Nets" and the wires translate to neural connections. Each Cell-Net is a pre-trained DNN that models the Boolean-logic behavior of the corresponding standard cell. In the PE-Net, every neural connection is associated with a bias that represents a perturbation in the signal propagated by that connection. We utilize a recently proposed misclassification-driven training algorithm to sensitize and identify biases that are critical to the functioning of the accelerator for a given application workload. The proposed framework achieves up to 100% accuracy in fault-criticality classification in 16-bit and 32-bit PEs by using the criticality knowledge of only 2% of the total faults in a PE.
Arjun Chaudhuri, Ching-Yuan Chen, Jonti Talukdar, Siddarth Madala, Abhishek Kumar Dubey, Krishnendu Chakrabarty
ITC3
2021 A BIST-based Dynamic Obfuscation Scheme for Resilience against Removal and Oracle-guided Attacks*
abstract
BISTLock is a recently proposed logic-locking technique that integrates a barrier finite-state-machine (FSM) with the built-in self-test (BIST) controller. We demonstrate the vulnerability of BISTLock to removal/bypass attacks and develop countermeasures to make it resilient against not only removal attacks but any form of Oracle-guided attack. Removal resilience is achieved through the incorporation of an input-signal scrambler. We demonstrate the vulnerability of the standalone scrambler to the SAT attack and present a reconfigurable LFSR-based dynamic authenticator that achieves SAT resilience. The proposed solution provides dynamic obfuscation upon the application of an incorrect key and prevents Oracle access to the attacker. We also present a security analysis of the overall system against Oraclefree attacks such as BMC-based sequential SAT and the FSM reverse engineering attack. We evaluate the security strength of the proposed solution and show that hardware overhead is low for a broad set of benchmark circuits.
Jonti Talukdar, Amitabh Das, Sohrab Aftabjahani, Peilin Song, Krishnendu Chakrabarty
ITC1
2020 Functional Criticality Classification of Structural Faults in AI Accelerators
abstract
The ubiquitous application of deep neural networks (DNNs) has led to a rise in demand for artificial intelligence (AI) accelerators. This paper studies the problem of classifying structural faults in such an accelerator based on their functional criticality. We analyze the impact of stuck-at faults in the processing elements (PEs) of a $128 \times 128$ systolic array designed to perform classification on the MNIST dataset using both 32-bit and 16-bit data paths. We present a two-tier machine-learning (ML) based method to assess the functional criticality of these faults. We address the problem of minimizing misclassification by utilizing generative adversarial networks (GANs). The two-tier ML/GAN-based criticality assessment method leads to less than 1% test escapes during functional criticality evaluation.
Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty
ITC2