Benjamin Tan 0001

dblp:195/3070 · DBLP profile ↗
← Back
49ranked-venue papers
4as first author
45since 2021 · last 2026
0000-0002-7642-3638ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 4 first-author · 35 since 2021Security and privacy · 7 · 7 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PALM: Program Analysis and LLM Methods for Crafting SystemVerilog Assertions
abstract
A promising approach for security verification of a Register-Transfer Level (RTL) design is assertion-based verification (ABV), where desired properties are expressed as SystemVerilog Assertions (SVAs). To create assertions, verification engineers typically start with identifying the relevant modules and necessary variables that are relevant to a given property and then construct the assertion based on those variables. While there have been several attempts to automate assertion creation, prior work identified that automatically recognizing relevant modules and subsequently extracting the required variables within the found module to construct an SVA is a bottleneck. Recently, Large Language Models (LLMs) have emerged, demonstrating promising code generation capabilities. However, their application in helping to automate valid SVA generation, along with the combination of static analysis methods, remains not well explored. This work investigates whether, and to what extent, LLMs can assist in each stage of the automation pipeline or whether their promise requires more evidence to substantiate. This study identifies specific areas where Large Language Models (LLMs) yield measurable and practical improvements in a hybrid workflow, as well as areas where their limitations are evident.
Raheel Afsharmazayejani, Benjamin Tan 0001
DATE2
2026 SVQL: SystemVerilog Query Language
Nicholas Allison, Benjamin Tan 0001
ACM Great Lakes Symposium on VLSI2
2026 Behavioral Analysis of AI Code Generation Agents: Edit, Rewrite, and Repetition
abstract
Artificial intelligence code generation agents have become transformative tools in modern software development, yet their behavioral patterns remain poorly understood. This paper presents a study analyzing pull request patches from the AIDev dataset to characterize the behavioral signatures of five code generation agents (Claude Code, Copilot, Cursor, Devin, and OpenAI Codex) across the top five programming languages (TypeScript, Python, Go, Java, and C#) in the dataset. We investigate two key research questions: “Do agents edit or rewrite existing code?”, and “How repetitive is each agent’s generated code?” Using token-level similarity metrics (Jaccard, TF-IDF, and fuzzy matching) and repetition analysis (n-gram distributions and Shannon entropy), we characterize edit-rewrite behavior by whether new code closely resembles or substantially differs from existing code. Our results show that Claude Code tends toward lower-similarity changes and higher token diversity, Devin tends toward higher-similarity changes indicative of more incremental modification, and OpenAI Codex exhibits mixed patterns across similarity measures. These behavioral patterns provide insights into how different AI agents approach code generation tasks.
Mahdieh Abazar, Reyhaneh Farahmand, Gouri Ginde, Benjamin Tan 0001, Lorenzo De Carli
MSR4
2026 Runtime Fault Localization in Deep Neural Network Accelerators
abstract
Systolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Although fault detection and repair techniques have been proposed to enhance the robustness of systolic arrays, fault localization remains an open problem. We propose a fault tolerance framework including run-time based fault detection and fault localization, both leveraging functional data to generate checksums on-the-fly. This approach enables error detection and localization during normal operation, avoiding the need for dedicated test patterns or additional downtime. Experimental evaluation shows that the proposed fault localization architecture incurs an area overhead less than 2% for a 256× 256 systolic array. In simulations, the proposed method achieves 100% fault detection and localization in a 256× 256 systolic array.
Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty
ACM Trans. Design Autom. Electr. Syst.3
2025 Identifying System-on-Chip Security Assets with Structure-Based Analysis
abstract
In the evolving field of hardware design, ensuring the security of System-on-Chips (SoCs) has become increasingly vital. As SoCs grow in complexity, integrating components from various sources, the identification and protection of security assets are crucial to prevent vulnerabilities. Traditional methods of identifying these assets are manual and time-intensive. To address this challenge, automated tools for security asset identification are essential, enabling faster and more accurate detection of critical assets early in the design process. In this paper, we propose a framework for the automated identification of security assets within SoCs. By transforming register-transfer level (RTL) code into graphs and leveraging deep neural networks (DNNs) to classify assets based on their structural patterns, our approach can effectively differentiate between security and non-security assets. Experimental results show that the proposed method achieves high classification accuracy, with the model reaching up to 99% accuracy in identifying security assets, significantly reducing the need for manual intervention.
Wei-Kai Liu, Benjamin Tan 0001, Krishnendu Chakrabarty
DAC2
2025 Hiding in Plain Sight: On the Robustness of AI-Generated Code Detection
Saman Pordanesh, Sufiyan Bukhari, Benjamin Tan 0001, Lorenzo De Carli
DIMVA (2)3
2025 Scaling Attacks on Large Logic-Locked Designs
abstract
Researchers have developed numerous strategies to alleviate the threat of malicious third-party foundries, including logic locking and its numerous sophisticated variants for hardware intellectual property (IP) protection. Recent work at the register-transfer level has opened the door to “large-scale” locking of large IPs (comprising thousands of gates) with hundreds to thousands of key bits. Recent security evaluation of such techniques treats the locked design as a monolith and has suggested that large logic-locked designs are practically secure, even from powerful SAT-based attacks. In this work, we challenge such findings by proposing and evaluating a novel algorithmic method to de-obfuscate large logic-locked circuits by attacking a set of small sub-circuit cones. The algorithm chooses a sub-optimal set of sub-circuit cones and proposes an attack sequence on these cones by leveraging the observation that each locking key-bit is distributed across multiple sub-circuit cones of varying sizes. This Divide And Conquer SAT (DACSAT) attack framework can de-obfuscate large designs, like an AES IP comprising 300,000 gates, logic-locked with up to 50,000 keys in around 3600 seconds, while an out-of-the-box, state-of-the-art SAT attack tool fails.
Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Ramesh Karri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Interpretable CNN-Based Lithographic Hotspot Detection Through Error Marker Learning
abstract
As the technology node develops toward its physical limit, lithographic hotspot detection has become increasingly important and ever-challenging in the computer-aided design (CAD) flow. In recent years, convolutional neural networks (CNNs) have achieved great success in hotspot detection. However, the interpretability of their hotspot prediction has yet to be considered. Compared with conventional lithography simulation and pattern matching-based methods, the black-box nature of CNNs wavers their practical applications with confidence. In this article, we propose the first interpretable CNN-based hotspot detector capable of providing high-detection accuracy and reliable explanations for hotspot identification. Specifically, we augment the training dataset with expanded error markers obtained and preprocessed from lithography simulation, which are then learned by an encoder-decoder architecture as intermediate features. We additionally introduce coordinate attention in the encoder to facilitate better-feature extraction. By learning these error markers and part of their surrounding metals as root cause hotspot features, our architecture achieves the highest-hotspot accuracy of 99.78% and the lowest-false positive rate of 5.29% compared to all prior work. Moreover, our method demonstrates the best visual and quantitative interpretability results when applying CNN interpretation methods.
Xun Ye, Dan Feng 0001, Benjamin Tan 0001, Yuzhe Ma, Kang Liu 0017
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 FLAG: Finding Line Anomalies (in RTL code) with Generative AI
abstract
Bug detection in Hardware Design Languages (HDLs) is an important problem in the System-on-Chip (SoC) development cycle. It is crucial to find defects at the earliest stage possible. While most fault localization requires the use of “tests” (e.g., test benches, fuzzing, and assertions) and a simulation or emulation framework, the advent of Large Language Models (LLMs) provides an opportunity for a test-free fault localization approach. This article proposes such a tool, called FLAG, which can identify functional and security defects in Register Transfer Level (RTL) code without synthesis or simulation. FLAG combines syntactic and generative AI techniques to implement fault localization in RTL code. It takes an RTL design as an input and outputs a set of line(s) that likely contain defects. It targets elements of RTL code most likely to contain bugs through static analysis means and then implements token-level and line-level analysis to obtain differences in original code and code generated by LLM to identify a line as buggy or not. The token-level approach evaluates each generated token (one at a time) and the line level approach evaluates the entire line generated by the LLM. We evaluate our approach on a corpus of synthetic and real-world bugs, of both functional and security related issues, in Verilog and SystemVerilog. Using line-level analysis, FLAG can identify 38 out of 120 real-world bugs and using token-level analysis, FLAG can identify 32 out of 81 synthetic bugs through the top-5 most likely bug locations identified without tests.
Baleegh Ahmad, Joey Ah-kiow, Benjamin Tan 0001, Ramesh Karri, Hammond A. Pearce
ACM Trans. Design Autom. Electr. Syst.3
2025 Automatically Improving LLM-based Verilog Generation using EDA Tool Feedback
abstract
Traditionally, digital hardware designs are written in the Verilog hardware description language (HDL) and debugged manually by engineers. This can be time-consuming and error-prone for complex designs. Large Language Models (LLMs) are emerging as a potential tool to help generate fully functioning HDL code, but most works have focused on generation in the single-shot capacity: i.e., run and evaluate, a process that does not leverage debugging and, as such, does not adequately reflect a realistic development process. In this work, we evaluate the ability of LLMs to leverage feedback from electronic design automation (EDA) tools to fix mistakes in their own generated Verilog. To accomplish this, we present an open-source, highly customizable framework, AutoChip, which combines conversational LLMs with the output from Verilog compilers and simulations to iteratively generate and repair Verilog. To determine the success of these LLMs we leverage the VerilogEval benchmark set. We evaluate four state-of-the-art conversational LLMs, focusing on readily accessible commercial models. EDA tool feedback proved to be consistently more effective than zero-shot prompting only with GPT-4o, the most computationally complex model we evaluated. In the best case, we observed a 5.8% increase in the number of successful designs with a 34.2% decrease in cost over the best zero-shot results. Mixing smaller models with this larger model at the end of the feedback iterations resulted in equally as much success as with GPT-4o using feedback, but incurred 41.9% lower cost (corresponding to an overall decrease in cost over zero-shot by 89.6%).
Jason Blocklove, Shailja Thakur, Benjamin Tan 0001, Hammond A. Pearce, Siddharth Garg, Ramesh Karri
ACM Trans. Design Autom. Electr. Syst.3
2025 ARIANNA: An Automatic Design Flow for Fabric Customization and eFPGA Redaction
abstract
In the modern global Integrated Circuit (IC) supply chain, protecting intellectual property (IP) is a complex challenge, and balancing IP loss risk and added cost for theft countermeasures is hard to achieve. Using embedded configurable logic allows designers to completely hide the functionality of selected design portions from parties that do not have access to the configuration string (bitstream). However, the design space of redacted solutions is huge, with tradeoffs between the portions selected for redaction and the configuration of the configurable embedded logic. We propose ARIANNA, a complete flow that aids the designer in all the stages, from selecting the logic to be hidden to tailoring the bespoke fabrics for the configurable logic used to hide it. We present a security evaluation of the considered fabrics and introduce two heuristics for the novel bespoke fabric flow. We evaluate the heuristics against an exhaustive approach. We also evaluate the complete flow using a selection of benchmarks. Results show that using ARIANNA to customize the redaction fabrics yields up to 3.3× lower overheads and 4× higher eFPGA fabric utilization than a one-fits-all fabric as proposed in prior works.
Luca Collini, Jitendra Bhandari, Chiara Muscari Tomajoli, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Xifan Tang, Pierre-Emmanuel Gaillardon, Ramesh Karri, Christian Pilato
ACM Trans. Design Autom. Electr. Syst.5
2025 LithoExp: Explainable Two-stage CNN-based Lithographic Hotspot Detection with Layout Defect Localization
abstract
Convolutional neural networks (CNNs) successfully detect lithographic hotspots by learning from hand-designed features of layout patterns or entire layouts, as images, in an end-to-end fashion. However, compared to lithography simulation, CNN-based solutions demonstrate inferior hotspot detection accuracy and a high false-alarm rate. Moreover, the interpretability of the hotspot prediction process has yet to be considered due to the “black-box” nature of CNNs. In this work, inspired by conventional lithography simulation where defect regions are simulated as direct evidence for hotspot identification, we propose an explainable two-stage CNN-based hotspot detector that considers both the accuracy and interpretability of hotspot detection. Our architecture learns to locate the defect areas in the first stage as extracted hotspot features. In the second stage, we combine the strength of feature engineering and end-to-end learning, incorporating the original layout input, the learned defect location map from the first stage, and a fixed auxiliary region of interest (ROI) map for final hotspot detection. Experimental results for our technique exhibit the highest hotspot accuracy (98.1%) and the lowest false-alarm rate (4.0%) thus far compared to all prior CNN solutions. We also demonstrate the best overall qualitative and quantitative interpretability results with the highest increase in confidence (IC) and the lowest average drop (AD) in scores when CNN interpretation methods such as Grad-CAM-based approaches are applied. We further demonstrate use cases of our technique for successfully justifying and pinpointing hotspot mispredictions by examining the prediction evidence from our learned defect locations.
Dan Feng 0001, Zhiyao Xie, Benjamin Tan 0001, Kang Liu 0017
ACM Trans. Design Autom. Electr. Syst.5
2025 Patchability-Driven Design Exploration for System-on-Chip Patching Architectures
abstract
As System-on-Chip (SoC) designs become increasingly complex, ensuring comprehensive verification has become more challenging, leading to overlooked hardware bugs that can be found in the field. Addressing hardware bugs post-deployment is difficult, as they typically cannot be easily fixed like software bugs. To tackle this issue, hardware-based patching mechanisms have emerged as a potential solution for providing in-field fixes. However, the lack of a standardized method to evaluate the ”patchability” of different designs complicates the integration of patching infrastructure into SoCs. In this article, we propose a fully parameterized Patch Support Block (PSB) architecture that can be tailored for various hardware designs, enabling post-deployment patching. We introduce a novel patchability score formulation that provides a quantifiable metric for evaluating the effectiveness of patching designs. Our approach considers both the observability and controllability of the patching hardware and provides a framework for system integrators to maximize patchability while managing resource constraints. Through experimentation with multiple design configurations, we demonstrate how our methodology can enhance patchability in hardware systems and provide security-related fixes for SoCs in real-world scenarios.
Wei-Kai Liu, Benjamin Tan 0001, Krishnendu Chakrabarty
ACM Trans. Design Autom. Electr. Syst.2
2024 Theoretical Patchability Quantification for IP-Level Hardware Patching Designs
abstract
As the complexity of System-on-Chip (SoC) designs continues to increase, ensuring thorough verification becomes a significant challenge for system integrators. The complexity of verification can result in undetected bugs. Unlike software or firmware bugs, hardware bugs are hard to fix after deployment and they require additional logic, i.e., patching logic integrated with the design in advance in order to patch. However, the absence of a standardized metric for defining “patchability” leaves system integrators relying on their understanding of each IP and security requirements to engineer ad hoc patching designs. In this paper, we propose a theoretical patchability quantification method to analyze designs at the Register Transfer Level (RTL) with provided patching options. Our quantification defines patchability as a combination of observability and controllability so that we can analyze and compare the patchability of IP variations. This quantification is a systematic approach to estimate each patching architecture’s ability to patch at run-time and complements existing patching works. In experiments, we compare several design options of the same patching architecture and discuss their differences in terms of theoretical patchability and how many potential weaknesses can be mitigated.
Wei-Kai Liu, Benjamin Tan 0001, Jason M. Fung, Krishnendu Chakrabarty
ASPDAC2
2024 Effective Runtime Fault Detection for DNN Accelerators
abstract
Systolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing of matrix multiplication, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Even though many algorithm-based fault tolerance (ABFT) algorithms have been proposed to detect and correct errors in matrix multiplication, these ABFT methods cannot detect many errors originating from the accelerator hardware. We propose a run-time based fault detection technique leveraging functional data to generate checksums on-the-fly, avoiding the requirement for test patterns. Experimental evaluation shows that the proposed fault detection architecture can achieve 100% test coverage while incurring an area overhead of less than 2% for a 256 × 256 systolic array.
Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty
ATS3
2024 Retrieval-Guided Reinforcement Learning for Boolean Circuit Minimization
abstract
Logic synthesis, a pivotal stage in chip design, entails optimizing chip specifications encoded in hardware description languages like Verilog into highly efficient implementations using Boolean logic gates. The process involves a sequential application of logic minimization heuristics (``synthesis recipe"), with their arrangement significantly impacting crucial metrics such as area and delay. Addressing the challenge posed by the broad spectrum of hardware design complexities — from variations of past designs (e.g., adders and multipliers) to entirely novel configurations (e.g., innovative processor instructions) — requires a nuanced 'synthesis recipe' guided by human expertise and intuition. This study conducts a thorough examination of learning and search techniques for logic synthesis, unearthing a surprising revelation: pre-trained agents, when confronted with entirely novel designs, may veer off course, detrimentally affecting the search trajectory. We present ABC-RL, a meticulously tuned $\alpha$ parameter that adeptly adjusts recommendations from pre-trained agents during the search process. Computed based on similarity scores through nearest neighbor retrieval from the training dataset, ABC-RL yields superior synthesis recipes tailored for a wide array of hardware designs. Our findings showcase substantial enhancements in the Quality of Result (QoR) of synthesized circuits, boasting improvements of up to 24.8\% compared to state-of-the-art techniques. Furthermore, ABC-RL achieves an impressive up to 9x reduction in runtime (iso-QoR) when compared to current state-of-the-art methodologies.
Animesh Basak Chowdhury, Marco Romanelli 0002, Benjamin Tan 0001, Ramesh Karri, Siddharth Garg
ICLR3
2024 Investigating the Feasibility of eFPGA-based Hardware Patching
abstract
System-on-Chip (SoC) designs are becoming increasingly complex, with the ability to detect and address all possible bugs at design time is highly challenging. Thus, to improve the survivability of SoC designs, it is desirable to be able to patch newly discovered design bugs or potential vulnerabilities in the field. Recently, the idea of hardware based patching, especially of hardware bugs, has emerged as a complementary approach to software/firmware-based post deployment updates. In anticipating potential problems, designers must invest an upfront cost to implement hardware-based patching infrastructures. In this paper, We investigate the feasibility of incorporating an embedded field-programmable gate array (eFPGA) fabric as an approach to enable hardware-based patching, i.e., reprogrammable hardware to patch hardware bugs. We propose, discuss, and evaluate three integration design architectures, characterizing the potential area and performance costs for each patching architecture and providing insights into how such architectures might be used to patch hardware bugs. Through a case study on an OpenPiton-based SoC, our results show the architectures’ area overhead costs ranging from $4.78 \%$ to $56.17 \%$, with latency arising from our example patches ranging from 1 to 2 cycles.
Anudeep Dharavathu, Benjamin Tan 0001
IOLTS2
2024 Extending ISO 15118-20 EV Charging: Preventing Downgrade Attacks and Enabling New Security Capabilities
abstract
Previous works have identified that EV charging can be weaponised to attack the power grid. As a case study, we consider the newest charging protocol ISO 15118–20, which provides a high-level communication protocol for EV charging. We first highlight fundamental issues in ISO 15118–20 which prevent the development of security features within the existing standard: We show that an attacker can perform a downgrade attack on ISO 15118–20, and propose modifications to the standard to prevent this. We show how this can be used to enable the development of additional security features within the modified protocol. A proof of concept is developed to prove functionality, determine interoperability between various parties, verify that it meets the original standard's timing requirements, and does not impact charging speed nor noticeably affect the length of a charging session.
Ross Porter, Morteza Biglari-Abhari, Benjamin Tan 0001, Duleepa J. Thrimawithana
PST3
2024 On Hardware Security Bug Code Fixes by Prompting Large Language Models
abstract
Novel AI-based code-writing Large Language Models (LLMs) such as OpenAI’s Codex have demonstrated capabilities in many coding-adjacent domains. In this work, we consider how LLMs may be leveraged to automatically repair identified security-relevant bugs present in hardware designs by generating replacement code. We focus on bug repair in code written in Verilog. For this study, we curate a corpus of domain-representative hardware security bugs. We then design and implement a framework to quantitatively evaluate the performance of any LLM tasked with fixing the specified bugs. The framework supports design space exploration of prompts (i.e., prompt engineering) and identifying the best parameters for the LLM. We show that an ensemble of LLMs can repair all fifteen of our benchmarks. This ensemble outperforms a state-of-the-art automated hardware bug repair tool on its own suite of bugs. These results show that LLMs have the ability to repair hardware security bugs and the framework is an important step towards the ultimate goal of an automated end-to-end bug repair tool.
Baleegh Ahmad, Shailja Thakur, Benjamin Tan 0001, Ramesh Karri, Hammond A. Pearce
IEEE Trans. Inf. Forensics Secur.3
2024 (Security) Assertions by Large Language Models
abstract
The security of computer systems typically relies on a hardware root of trust. As vulnerabilities in hardware can have severe implications on a system, there is a need for techniques to support security verification activities. Assertion-based verification is a popular verification technique that involves capturing design intent in a set of assertions that can be used in formal verification or testing-based checking. However, writing security-centric assertions is a challenging task. In this work, we investigate the use of emerging large language models (LLMs) for code generation in hardware assertion generation for security, where primarily natural language prompts, such as those one would see as code comments in assertion files, are used to produce SystemVerilog assertions. We focus our attention on a popular LLM and characterize its ability to write assertions out of the box, given varying levels of detail in the prompt. We design an evaluation framework that generates a variety of prompts, and we create a benchmark suite comprising real-world hardware designs and corresponding golden reference assertions that we want to generate with the LLM.
Rahul Kande, Hammond A. Pearce, Benjamin Tan 0001, Brendan Dolan-Gavitt, Shailja Thakur, Ramesh Karri, Jeyavijayan Rajendran
IEEE Trans. Inf. Forensics Secur.3
2024 VeriGen: A Large Language Model for Verilog Code Generation
abstract
In this study, we explore the capability of Large Language Models (LLMs) to automate hardware design by automatically completing partial Verilog code, a common language for designing and modeling digital systems. We fine-tune pre-existing LLMs on Verilog datasets compiled from GitHub and Verilog textbooks. We evaluate the functional correctness of the generated Verilog code using a specially designed test suite, featuring a custom problem set and testing benches. Here, our fine-tuned open-source CodeGen-16B model outperforms the commercial state-of-the-art GPT-3.5-turbo model with a 1.1% overall increase. Upon testing with a more diverse and complex problem set, we find that the fine-tuned model shows competitive performance against state-of-the-art gpt-3.5-turbo, excelling in certain scenarios. Notably, it demonstrates a 41% improvement in generating syntactically correct Verilog code across various problem categories compared to its pre-trained counterpart, highlighting the potential of smaller, in-house LLMs in hardware design automation. We release our training/evaluation scripts and LLM checkpoints as open-source contributions.
Shailja Thakur, Baleegh Ahmad, Hammond A. Pearce, Benjamin Tan 0001, Brendan Dolan-Gavitt, Ramesh Karri, Siddharth Garg
ACM Trans. Design Autom. Electr. Syst.4
2023 ALMOST: Adversarial Learning to Mitigate Oracle-less ML Attacks via Synthesis Tuning
abstract
Oracle-less machine learning (ML) attacks have broken various logic locking schemes. Regular synthesis, which is tailored for area-power-delay optimization, yields netlists where key-gate localities are vulnerable to learning. Thus, we call for security-aware logic synthesis. We propose ALMOST, a framework for adversarial learning to mitigate oracle-less ML attacks via synthesis tuning. ALMOST uses a simulated-annealing-based synthesis recipe generator, employing adversarially trained models that can predict state-of-the-art attacks’ accuracies over wide ranges of recipes and key-gate localities. Experiments on ISCAS benchmarks confirm the attacks’ accuracies drops to around 50% for ALMOST-synthesized circuits, all while not undermining design optimization.
Animesh Basak Chowdhury, Lilas Alrahis, Luca Collini, Johann Knechtel, Ramesh Karri, Siddharth Garg, Ozgur Sinanoglu, Benjamin Tan 0001
DAC8
2023 Benchmarking Large Language Models for Automated Verilog RTL Code Generation
abstract
Automating hardware design could obviate a signif-icant amount of human error from the engineering process and lead to fewer errors. Verilog is a popular hardware description language to model and design digital systems, thus generating Verilog code is a critical first step. Emerging large language models (LLMs) are able to write high-quality code in other programming languages. In this paper, we characterize the ability of LLMs to generate useful Verilog. For this, we fine-tune pre-trained LLMs on Verilog datasets collected from GitHub and Verilog textbooks. We construct an evaluation framework comprising test-benches for functional analysis and a flow to test the syntax of Verilog code generated in response to problems of varying difficulty. Our findings show that across our problem scenarios, the fine-tuning results in LLMs more capable of producing syntactically correct code (25.9% overall). Further, when analyzing functional correctness, a fine-tuned open-source CodeGen LLM can outperform the state-of-the-art commercial Codex LLM (6.5% overall). We release our training/evaluation scripts and LLM checkpoints as open source contributions.
Shailja Thakur, Baleegh Ahmad, Zhenxing Fan, Hammond A. Pearce, Benjamin Tan 0001, Ramesh Karri, Brendan Dolan-Gavitt, Siddharth Garg
DATE5
2023 Examining Zero-Shot Vulnerability Repair with Large Language Models
abstract
Human developers can produce code with cybersecurity bugs. Can emerging ‘smart’ code completion tools help repair those bugs? In this work, we examine the use of large language models (LLMs) for code (such as OpenAI’s Codex and AI21’s Jurassic J-1) for zero-shot vulnerability repair. We investigate challenges in the design of prompts that coax LLMs into generating repaired versions of insecure code. This is difficult due to the numerous ways to phrase key information— both semantically and syntactically—with natural languages. We perform a large scale study of five commercially available, black-box, "off-the-shelf" LLMs, as well as an open-source model and our own locally-trained model, on a mix of synthetic, hand-crafted, and real-world security bug scenarios. Our experiments demonstrate that while the approach has promise (the LLMs could collectively repair 100% of our synthetically generated and hand-crafted scenarios), a qualitative evaluation of the model’s performance over a corpus of historical real-world examples highlights challenges in generating functionally correct code.
Hammond A. Pearce, Benjamin Tan 0001, Baleegh Ahmad, Ramesh Karri, Brendan Dolan-Gavitt
SP2
2023 Examining Zero-Shot Vulnerability Repair with Large Language Models
abstract
Human developers can produce code with cybersecurity bugs. Can emerging ‘smart’ code completion tools help repair those bugs? In this work, we examine the use of large language models (LLMs) for code (such as OpenAI’s Codex and AI21’s Jurassic J-1) for zero-shot vulnerability repair. We investigate challenges in the design of prompts that coax LLMs into generating repaired versions of insecure code. This is difficult due to the numerous ways to phrase key information— both semantically and syntactically—with natural languages. We perform a large scale study of five commercially available, black-box, "off-the-shelf" LLMs, as well as an open-source model and our own locally-trained model, on a mix of synthetic, hand-crafted, and real-world security bug scenarios. Our experiments demonstrate that while the approach has promise (the LLMs could collectively repair 100% of our synthetically generated and hand-crafted scenarios), a qualitative evaluation of the model’s performance over a corpus of historical real-world examples highlights challenges in generating functionally correct code.
Hammond A. Pearce, Benjamin Tan 0001, Baleegh Ahmad, Ramesh Karri, Brendan Dolan-Gavitt
SP2
2023 Bulls-Eye: Active Few-Shot Learning Guided Logic Synthesis
abstract
Generating suboptimal synthesis transformation sequences (“synthesis recipe”) is an important problem in logic synthesis. Manually crafted synthesis recipes have poor quality. State-of-the art machine learning (ML) works to generate synthesis recipes do not scale to large netlists as the models need to be trained from scratch, for which training data is collected using time-consuming synthesis runs. We propose a new approach, Bulls-Eye, that fine-tunes a pretrained model on past synthesis data to accurately predict the quality of a synthesis recipe for an unseen netlist. Our approach achieves$2\times $–$30\times $runtime improvement and generates synthesis recipes achieving close to 95% quality-of-result (QoR) compared to conventional techniques using actual synthesis runs. We show our QoR beat state-of-the-art approaches on various benchmarks.
Animesh Basak Chowdhury, Benjamin Tan 0001, Ryan Carey, Tushit Jain, Ramesh Karri, Siddharth Garg
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Hardware-Supported Patching of Security Bugs in Hardware IP Blocks
abstract
To satisfy various design requirements and application needs, designers integrate multiple intellectual property blocks (IPs) to produce a system on chip (SoC). For improved survivability, designers should be able to patch the SoC to mitigate potential security issues arising from hardware IPs; for increased flexibility, we propose adding programmable hardware-based support for monitoring and bug mitigation. However, it is a challenge to decide how much additional cost a designer should expend up front to deal with unknown, future issues. We propose an approach that guides designers toward maximizing the benefits of adding “patchability” to various IPs in the system, given a target resource overhead. We frame the design problem as an integer quadratic program and show that our approach achieves superior patchability compared to the naïve and baseline approaches for a given cost limit. Experimental results show that when we set a cost limit of 2% field-programmable gate array adaptive logic module usage, our solution can generate a viable patching infrastructure with six patching blocks offering patches for seven different services in our case study.
Wei-Kai Liu, Benjamin Tan 0001, Jason M. Fung, Ramesh Karri, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 High-Level Approaches to Hardware Security: A Tutorial
abstract
Designers use third-party intellectual property (IP) cores and outsource various steps in the integrated circuit (IC) design and manufacturing flow. As a result, security vulnerabilities have been rising. This is forcing IC designers and end users to re-evaluate their trust in ICs. If attackers get hold of an unprotected IC, they can reverse engineer the IC and pirate the IP. Similarly, if attackers get hold of a design, they can insert malicious circuits or take advantage of “backdoors” in a design. Unintended design bugs can also result in security weaknesses. This tutorial paper provides an introduction to the domain of hardware security through two pedagogical examples of hardware security problems. The first is a walk-through of the scan chain-based side channel attack. The second is a walk-through of logic locking of digital designs. The tutorial material is accompanied by open access digital resources that are linked in this article.
Hammond A. Pearce, Ramesh Karri, Benjamin Tan 0001
ACM Trans. Embed. Comput. Syst.3
2023 Not All Fabrics Are Created Equal: Exploring eFPGA Parameters for IP Redaction
abstract
Semiconductor design houses rely on third-party foundries to manufacture their integrated circuits (ICs). While this trend allows them to tackle fabrication costs, it introduces security concerns as external (and potentially malicious) parties can access critical parts of the designs and steal or modify the intellectual property (IP). Embedded field-programmable gate array (eFPGA) redaction is a promising technique to protect critical IPs of an ASIC by redacting (i.e., removing) critical parts and mapping them onto a custom reconfigurable fabric. Only trusted parties will receive the correct bitstream to restore the redacted functionality. While previous studies imply that using an eFPGA is a sufficient condition to provide security against IP threats like reverse-engineering, whether this truly holds for all eFPGA architectures is unclear, thus motivating the study in this article. We examine the security of eFPGA fabrics generated by varying different FPGA design parameters. We characterize the power, performance, and area (PPA) characteristics and evaluate each fabric’s resistance to Boolean satisfiability (SAT)-based bitstream recovery. Our results encourage designers to work with custom eFPGA fabrics rather than off-the-shelf commercial FPGAs and reveals that only considering a redaction fabric’s bitstream size is inadequate for gauging security.
Jitendra Bhandari, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Christian Pilato, Ganesh Gore, Xifan Tang, Scott Temple, Pierre-Emmanuel Gaillardon, Ramesh Karri
IEEE Trans. Very Large Scale Integr. Syst.3
2022 High-level design methods for hardware security: is it the right choice? invited
abstract
Due to the globalization of the electronics supply chain, hardware engineers are increasingly interested in modifying their chip designs to protect their intellectual property (IP) or the privacy of the final users. However, the integration of state-of-the-art solutions for hardware and hardware-assisted security is not fully automated, requiring the amendment of stable tools and industrial toolchains. This significantly limits the application in industrial designs, potentially affecting the security of the resulting chips. We discuss how existing solutions can be adapted to implement security features at higher levels of abstractions (during high-level synthesis or directly at the register-transfer level) and complement current industrial design and verification flows. Our modular framework allows designers to compose these solutions and create additional protection layers.
Christian Pilato, Donatella Sciuto, Benjamin Tan 0001, Siddharth Garg, Ramesh Karri
DAC3
2022 Designing ML-resilient locking at register-transfer level
abstract
Various logic-locking schemes have been proposed to protect hardware from intellectual property piracy and malicious design modifications. Since traditional locking techniques are applied on the gate-level netlist after logic synthesis, they have no semantic knowledge of the design function. Data-driven, machine-learning (ML) attacks can uncover the design flaws within gate-level locking. Recent proposals on register-transfer level (RTL) locking have access to semantic hardware information. We investigate the resilience of ASSURE, a state-of-the-art RTL locking method, against ML attacks. We used the lessons learned to derive two ML-resilient RTL locking schemes built to reinforce ASSURE locking. We developed ML-driven security metrics to evaluate the schemes against an RTL adaptation of the state-of-the-art, ML-based SnapShot attack.
Dominik Germek, Luca Collini, Benjamin Tan 0001, Christian Pilato, Ramesh Karri, Rainer Leupers
DAC3
2022 ALICE: an automatic design flow for eFPGA redaction
abstract
Fabricating an integrated circuit is becoming unaffordable for many semiconductor design houses. Outsourcing the fabrication to a third-party foundry requires methods to protect the intellectual property of the hardware designs. Designers can rely on embedded reconfigurable devices to completely hide the real functionality of selected design portions unless the configuration string (bitstream) is provided. However, selecting such portions and creating the corresponding reconfigurable fabrics are still open problems. We propose ALICE, a design flow that addresses the EDA challenges of this problem. ALICE partitions the RTL modules between one or more reconfigurable fabrics and the rest of the circuit, automating the generation of the corresponding redacted design.
Chiara Muscari Tomajoli, Luca Collini, Jitendra Bhandari, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Xifan Tang, Pierre-Emmanuel Gaillardon, Ramesh Karri, Christian Pilato
DAC5
2022 Don't CWEAT It: Toward CWE Analysis Techniques in Early Stages of Hardware Design
abstract
To help prevent hardware security vulnerabilities from propagating to later design stages where fixes are costly, it is crucial to identify security concerns as early as possible, such as in RTL designs. In this work, we investigate the practical implications and feasibility of producing a set of security-specific scanners that operate on Verilog source files. The scanners indicate parts of code that might contain one of a set of MITRE's common weakness enumerations (CWEs). We explore the CWE database to characterize the scope and attributes of the CWEs and identify those that are amenable to static analysis. We prototype scanners and evaluate them on 11 open source designs - 4 system-on-chips (SoC) and 7 processor cores - and explore the nature of identified weaknesses. Our analysis reported 53 potential weaknesses in the OpenPiton SoC used in [email protected], 11 of which we confirmed as security concerns.
Baleegh Ahmad, Wei-Kai Liu, Luca Collini, Hammond A. Pearce, Jason M. Fung, Jonathan Valamehr, Mohammad Bidmeshki, Piotr Sapiecha, Krishnendu Chakrabarty, Ramesh Karri, Benjamin Tan 0001
ICCAD12
2022 Reconfigurable Logic for Hardware IP Protection: Opportunities and Challenges
abstract
Protecting the intellectual property (IP) of integrated circuit (IC) design is becoming a significant concern of fab-less semiconductor design houses. Malicious actors can access the chip design at any stage, reverse engineer the functionality, and create illegal copies. On the one hand, defenders are crafting more and more solutions to hide the critical portions of the circuit. On the other hand, attackers are designing more and more powerful tools to extract useful information from the design and reverse engineer the functionality, especially when they can get access to working chips. In this context, the use of custom reconfigurable fabrics has recently been investigated for hardware IP protection. This paper will discuss recent trends in hardware obfuscation with embedded FPGAs, focusing also on the open challenges that must be necessarily addressed for making this solution viable.
Luca Collini, Benjamin Tan 0001, Christian Pilato, Ramesh Karri
ICCAD2
2022 Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
abstract
There is burgeoning interest in designing AI-based systems to assist humans in designing computing systems, including tools that automatically generate computer code. The most notable of these comes in the form of the first self-described ‘AI pair programmer’, GitHub Copilot, which is a language model trained over open-source GitHub code. However, code often contains bugs—and so, given the vast quantity of unvetted code that Copilot has processed, it is certain that the language model will have learned from exploitable, buggy code. This raises concerns on the security of Copilot’s code contributions. In this work, we systematically investigate the prevalence and conditions that can cause GitHub Copilot to recommend insecure code. To perform this analysis we prompt Copilot to generate code in scenarios relevant to high-risk cybersecurity weaknesses, e.g. those from MITRE’s “Top 25” Common Weakness Enumeration (CWE) list. We explore Copilot’s performance on three distinct code generation axes—examining how it performs given diversity of weaknesses, diversity of prompts, and diversity of domains. In total, we produce 89 different scenarios for Copilot to complete, producing 1,689 programs. Of these, we found approximately 40% to be vulnerable.
Hammond A. Pearce, Baleegh Ahmad, Benjamin Tan 0001, Brendan Dolan-Gavitt, Ramesh Karri
SP3
2022 Innovation Practices Track: Security in Test and Test for Security
abstract
VLSI testing is essential to guarantee the correct functionality of the chip design. The recent advances in hardware security have posed new challenges for testing. In this IP session, we discuss the security in test and test for security through three talks. First, we give a brief overview of the security vulnerabilities and countermeasures in scan chain design, followed by a detailed discussion of a new configurable partial scan design approach. Second, we present the challenges in testing the security of design at various design stages and propose a strategy to identify potential security vulnerabilities in early design stages. Finally, we consider physical unclonable function (PUF) and develop an adaptive framework based on machine learning for the test and error correction of PUF designs.
Gang Qu 0001, Benjamin Tan 0001, Kuheli Pratihar, Debdeep Mukhopadhyay, Ramesh Karri
VTS2
2022 Robust Deep Learning for IC Test Problems
abstract
Numerous machine learning (ML), and more recently, deep-learning (DL)-based approaches, have been proposed to tackle scalability issues in electronic design automation, including those in integrated circuit (IC) test. This article examines state-of-the-art DL for IC test and highlights two critical unaddressed challenges. The first challenge involves identifying fit-for-purpose statistical metrics to train and evaluate ML model performance and usefulness in IC test. Our work shows that current metrics do not reflect how well ML models have learned to generalize and perform in the domain-specific context. From this insight, we propose and evaluate alternative metrics that better capture a model’s likely usefulness in the IC test problem. The second challenge is to choose an appropriate input abstraction so as to enable an ML model to learn robust and reliable features. We investigate how well DL for IC test techniques generalize by exploring their robustness to perturbations that alter a netlist’s structure but do not alter its functionality. This article provides insights into challenges via empirical evaluation of the state-of-the-art and offers guidance for future work.
Animesh Basak Chowdhury, Benjamin Tan 0001, Siddharth Garg, Ramesh Karri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Subverting Privacy-Preserving GANs: Hiding Secrets in Sanitized Images
abstract
Unprecedented data collection and sharing have exacerbated privacy concerns and led to increasing interest in privacy-preserving tools that remove sensitive attributes from images while maintaining useful information for other tasks. Currently, state-of-the-art approaches use privacy-preserving generative adversarial networks (PP-GANs) for this purpose, for instance, to enable reliable facial expression recognition without leaking users' identity. However, PP-GANs do not offer formal proofs of privacy and instead rely on experimentally measuring information leakage using classification accuracy on the sensitive attributes of deep learning (DL)-based discriminators. In this work, we question the rigor of such checks by subverting existing privacy-preserving GANs for facial expression recognition. We show that it is possible to hide the sensitive identification data in the sanitized output images of such PP-GANs for later extraction, which can even allow for reconstruction of the entire input images, while satisfying privacy checks. We demonstrate our approach via a PP-GAN-based architecture and provide qualitative and quantitative evaluations using two public datasets. Our experimental results raise fundamental questions about the need for more rigorous privacy checks of PP-GANs, and we provide insights into the social impact of these.
Kang Liu 0017, Benjamin Tan 0001, Siddharth Garg
AAAI2
2021 Attacking a CNN-based Layout Hotspot Detector Using Group Gradient Method
abstract
Deep neural networks are being used in disparate VLSI design automation tasks, including layout printability estimation, mask optimization, and routing congestion analysis. Preliminary results show the power of deep learning as an alternate solution in state-of-the-art design and sign-off flows. However, deep learning is vulnerable to adversarial attacks. In this paper, we examine the risk of state-of-the-art deep learning-based layout hotspot detectors under practical attack scenarios. We show that legacy gradient-based attacks do not adequately consider the design rule constraints. We present an innovative adversarial attack formulation to attack the layout clips and propose a fast group gradient method to solve it. Experiments show that the attack can deceive the deep neural networks using small perturbations in clips which preserve layout functionality while meeting the design rules. The source code is available at https://github.com/phdyang007/dlhsd/tree/dct_as_conv.
Shifan Zhang, Kang Liu 0017, Siting Liu 0002, Benjamin Tan 0001, Ramesh Karri, Siddharth Garg, Bei Yu 0001, Evangeline F. Y. Young
ASP-DAC5
2021 Invited: Independent Verification and Validation of Security-Aware EDA Tools and IP
abstract
Secure silicon requires a seamless integration of new tools, new IP, and design flows to help designers protect integrated circuits from increasingly sophisticated attacks. Independent Validation and Verification (IV&V) of this integrated technology is important to ensure that the tools actually deliver on their security claims when used by independent parties (i.e., people who were not involved in designing the tools). This work discusses the principles and approaches for IV&V of such a complex design environment, including validation of the security strength of the various hardware security techniques, such as combinational and sequential logic locking, Trojan Detection, side-channel mitigation, and blockchain-based asset management. The main challenge in running an IV&V effort is to ensure that the process provides rigorous, methodical and provable evaluation of the claims of not only the component tools and IP, but whether such an integrated environment can produce security-hardened designs by a non-security expert. CCS Concepts • Hardware $\rightarrow$ Very large scale integration design; Methodologies for EDA; • Security and privacy $\rightarrow$ Security in hardware.
Benjamin Tan 0001, Siddharth Garg, Ramesh Karri, Yuntao Liu 0001, Michael Zuzak, Abhisek Chakraborty, Ankur Srivastava 0001, Omid Aramoon, Qian Xu 0022, Gang Qu 0001, Adam A. Porter, Jeno Szep, Warren Savage
DAC1
2021 Exploring eFPGA-based Redaction for IP Protection
abstract
Recently, eFPGA-based redaction has been proposed as a promising solution for hiding parts of a digital design from untrusted entities, where legitimate end-users can restore functionality by loading the withheld bitstream after fabrication. However, when deciding which parts of a design to redact, there are a number of practical issues that designers need to consider, including area and timing overheads, as well as security factors. Adapting an open-source FPGA fabric generation flow, we perform a case study to explore the trade-offs when redacting different modules of open-source intellectual property blocks (IPs) and explore how different parts of an eFPGA contribute to the security. We provide new insights into the feasibility and challenges of using eFPGA-based redaction as a security solution.
Jitendra Bhandari, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Christian Pilato, Ganesh Gore, Xifan Tang, Scott Temple, Pierre-Emmanuel Gaillardon, Ramesh Karri
ICCAD3
2021 Special Session: Machine Learning for Semiconductor Test and Reliability
abstract
With technology scaling approaching atomic levels, IC test and diagnosis of complex System-on-Chips (SoCs) become overwhelming challenging. In addition, sustaining the reliability of transistors as well as circuits at such extreme feature sizes, for the entire projected lifetime, also become profoundly difficult. This holds even more when it comes to emerging technologies that go beyond convectional CMOS in which the underlying physics are not yet fully understood. In this special session paper, we describe the usage of machine learning in several test and reliability related areas. First, we demonstrate the vital role that machine learning can play in IC test showing the importance of explainability as a frontier for machine learning in IC test. Afterwards, we discuss how novel physics-informed neural networks can be employed to model electrostatic problems in VLSI designs. This is essential to mitigate the deleterious effects of of time dependent dielectric breakdown, which is the key source of reliability degradations. Finally, we discuss the major sources of reliability degradations at the transistor level in advanced technology nodes such as transistor aging phenomena and self-heating effects as well as we demonstrate how machine learning approaches can further help in developing reliable emerging technologies.
Hussam Amrouch, Animesh Basak Chowdhury, Wentian Jin, Ramesh Karri, Farshad Khorrami, Prashanth Krishnamurthy, Ilia Polian, Victor M. van Santen, Benjamin Tan 0001, Sheldon X.-D. Tan
VTS9
2021 Training Data Poisoning in ML-CAD: Backdooring DL-Based Lithographic Hotspot Detectors
abstract
Recent efforts to enhance computer-aided design (CAD) flows have seen the proliferation of machine learning (ML)-based techniques. However, despite achieving state-of-the-art performance in many domains, techniques, such as deep learning (DL) are susceptible to various adversarial attacks. In this work, we explore the threat posed by training data poisoning attacks where a malicious insider can try to insert backdoors into a deep neural network (DNN) used as part of the CAD flow. Using a case study on lithographic hotspot detection, we explore how an adversary can contaminate training data with specially crafted, yet meaningful, genuinely labeled, and design rule compliant poisoned clips. Our experiments show that very low poisoned/clean data ratio in training data is sufficient to backdoor the DNN; an adversary can “hide” specific hotspot clips at inference time by including a backdoor trigger shape in the input with ~100% success. This attack provides a novel way for adversaries to sabotage and disrupt the distributed design process. After finding that training data poisoning attacks are feasible and stealthy, we explore a potential ensemble defense against possible data contamination, showing promising attack success reduction. Our results raise fundamental questions about the robustness of DL-based systems in CAD, and we provide insights into the implications of these.
Kang Liu 0017, Benjamin Tan 0001, Ramesh Karri, Siddharth Garg
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Bias Busters: Robustifying DL-Based Lithographic Hotspot Detectors Against Backdooring Attacks
abstract
Deep learning (DL) offers potential improvements throughout the CAD tool-flow, one promising application being lithographic hotspot detection. However, DL techniques have been shown to be especially vulnerable to inference and training time adversarial attacks. Recent work has demonstrated that a small fraction of malicious physical designers can stealthily “backdoor” a DL-based hotspot detector during its training phase such that it accurately classifies regular layout clips but predicts hotspots containing a specially crafted trigger shape as nonhotspots. We propose a novel training data augmentation strategy as a powerful defense against such backdooring attacks. The defense works by eliminating the intentional biases introduced in the training data but does not require knowledge of which training samples are poisoned or the nature of the backdoor trigger. Our results show that the defense can drastically reduce the attack success rate from 84% to ~0%.
Kang Liu 0017, Benjamin Tan 0001, Gaurav Rajavendra Reddy, Siddharth Garg, Yiorgos Makris, Ramesh Karri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Toward Hardware-Based IP Vulnerability Detection and Post-Deployment Patching in Systems-on-Chip
abstract
System integrators create heterogeneous systems-on-chip (SoCs) by integrating numerous third-party intellectual property blocks (3PIPs) to achieve application-specific design goals. With increasing intellectual property (IP) complexity, 3PIPs can suffer from hardware bugs or they can inadvertently introduce other software-exploitable security threats to the SoC. To ensure the ongoing survivability of new SoCs, we need infrastructure for patching newly discovered IP issues after an SoC has been deployed. To address the increasing risks from 3PIPs, we explore the feasibility and limitations of implementing monitoring and mitigation capabilities in hardware. Our proposed monitoring and mitigation patch (MoP) blocks provide a defensive foundation against critical IP-centric issues, focusing on situations where a system integrator only has interface-level visibility of 3PIP designs. The MoPs are distributed throughout the SoC to monitor and mitigate issues directly in hardware and transparently for potentially compromised software-the MoPs are resilient against run-time compromised software and firmware. We ensure that these monitors are reconfigurable after deployment by implementing them using embedded-FPGAs or as a reprogrammable, fixed-design module. We perform a case study of numerous IP-types and model a selection of security-relevant issues and bugs in the IPs, exploring the relative complexity and potential resource overhead. Our study shows the utility of our proposed approach, with MoP blocks requiring less than ~1.5% of the adaptive logic modules (ALMs) in a Cyclone V FPGA for interface monitoring and issue mitigation per IP.
Benjamin Tan 0001, Rana Elnaggar, Jason M. Fung, Ramesh Karri, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Poisoning the (Data) Well in ML-Based CAD: A Case Study of Hiding Lithographic Hotspots
abstract
Machine learning (ML) provides state-of-the-art performance in many parts of computer-aided design (CAD) flows. However, deep neural networks (DNNs) are susceptible to various adversarial attacks, including data poisoning to compromise training to insert backdoors. Sensitivity to training data integrity presents a security vulnerability, especially in light of malicious insiders who want to cause targeted neural network misbehavior. In this study, we explore this threat in lithographic hotspot detection via training data poisoning, where hotspots in a layout clip can be "hidden" at inference time by including a trigger shape in the input. We show that training data poisoning attacks are feasible and stealthy, demonstrating a backdoored neural network that performs normally on clean inputs but misbehaves on inputs when a backdoor trigger is present. Furthermore, our results raise some fundamental questions about the robustness of ML-based systems in CAD.
Kang Liu 0017, Benjamin Tan 0001, Ramesh Karri, Siddharth Garg
DATE2
2020 Adversarial Perturbation Attacks on ML-based CAD: A Case Study on CNN-based Lithographic Hotspot Detection
abstract
There is substantial interest in the use of machine learning (ML)-based techniques throughout the electronic computer-aided design (CAD) flow, particularly those based on deep learning. However, while deep learning methods have surpassed state-of-the-art performance in several applications, they have exhibited intrinsic susceptibility to adversarial perturbations - small but deliberate alterations to the input of a neural network, precipitating incorrect predictions. In this article, we seek to investigate whether adversarial perturbations pose risks to ML-based CAD tools, and if so, how these risks can be mitigated. To this end, we use a motivating case study of lithographic hotspot detection, for which convolutional neural networks (CNN) have shown great promise. In this context, we show the first adversarial perturbation attacks on state-of-the-art CNN-based hotspot detectors; specifically, we show that small (on average 0.5% modified area), functionality preserving, and design-constraint-satisfying changes to a layout can nonetheless trick a CNN-based hotspot detector into predicting the modified layout as hotspot free (with up to 99.7% success in finding perturbations that flip a detector's output prediction, based on a given set of attack constraints). We propose an adversarial retraining strategy to improve the robustness of CNN-based hotspot detection and show that this strategy significantly improves robustness (by a factor of ∼3) against adversarial attacks without compromising classification accuracy.
Kang Liu 0017, Yuzhe Ma, Benjamin Tan 0001, Bei Yu 0001, Evangeline F. Y. Young, Ramesh Karri, Siddharth Garg
ACM Trans. Design Autom. Electr. Syst.4
2017 Towards decentralized system-level security for MPSoC-based embedded applications
Benjamin Tan 0001, Morteza Biglari-Abhari, Zoran A. Salcic
J. Syst. Archit.1
2017 An Automated Security-Aware Approach for Design of Embedded Systems on MPSoC
abstract
MPSoC-based embedded systems design is becoming increasingly complex. Not only do we need to satisfy multiple design objectives, we increasingly need to address potential security risks. In this work, we propose a security-aware systematic design approach which explores the design space, given a system-level application description, by generating potential architecture configurations of execution platform nodes that are interconnected using a NoC. We then perform automated security analysis to check the generated configurations against designer-specified security constraints. Following the analysis, we use an automated architecture configuration refinement process to generate a list of security additions that are inserted into the initial configuration so that the security constraints are satisfied. By performing this refinement on several candidate configuration options, we can explore the trade-off between resource cost and security. In this paper, we illustrate the proposed approach using a Smart Home Control System application.
Benjamin Tan 0001, Morteza Biglari-Abhari, Zoran A. Salcic
ACM Trans. Embed. Comput. Syst.1