Hammond A. Pearce

dblp:183/4902 · also Hammond Pearce · DBLP profile ↗
← Back
39ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0002-3488-7004ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 12 since 2021Security and privacy · 13 · 3 first-author · 13 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 7 since 2021Theory of computation · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SCREAM: Secure Channels for Real-time Evaluation of Additive Manufacturing
abstract
Additive Manufacturing (AM), also known as 3D printing, offers several advantages, including on-site production, enhanced throughput, and efficient use of raw materials. However, the rise in its usage has also led to an increase in potential threats that aim to disrupt the printing process. These attacks can subtly alter the design (CAD or STL) files or machine instructions (g-code), which can cause significant economic and reputational harm to the victim company. Current detection techniques, based on acoustic, magnetic, and accelerationbased side-channel analysis, have proven to be ineffective. Although power side-channel analysis is more effective than other means, it is expensive and not scalable. This paper proposes a novel detection method, SCREAM, that assumes the user has access to a trusted STL source and an untrusted g-code. SCREAM leverages the pulse trains sent to the motors to reconstruct the executing g-code. To ensure the safe and accurate execution of g-code, a three-level comparison is performed between recovered and untrusted g-code, as well as trusted STL ensuring successful detection of any anomalies present in the executing g-code. Our testing has shown that this method can detect a range of existing attacks on AM, including malicious firmware manipulation, FLAW3D, and Needle in a Haystack.
Prithwish Basu Roy, Jason Blocklove, Mudit Bhargava, Hammond A. Pearce, Prashanth Krishnamurthy, Ozgur Sinanoglu, Nikhil Gupta 0002, Farshad Khorrami, Ramesh Karri
AsiaCCS4
2026 ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
abstract
Achieving reproducibility, quantity, and diversity in vulnerability datasets has long been viewed as an inherent three-way trade-off, where improving one dimension often comes at the cost of the others. In practice, reproducibility has been the dimension most often neglected. This has limited what can be automatically extracted from historical bug datasets, and has reduced their utility for downstream security research. In this work, we propose a method to produce a new security dataset which ensures reproducibility for diverse vulnerabilities at scale by identifying the key obstacles to large-scale bug reproduction and addressing them with general solutions. Using this method, we introduce full reproducibility to the largest open source software vulnerability dataset (OSS-Fuzz) and construct the ARVO dataset (an Atlas of Reproducible Vulnerabilities in Open-source software). ARVO is a large-scale dataset consisting of over 6,100 real-world vulnerabilities across 311 projects. Focusing on reproducibility, ARVO differs from existing datasets by providing each vulnerability in a form that can be consistently rebuilt, triggered, and analyzed across versions. Reproducibility also enables automatic identification of the corresponding patch for each vulnerability and supports direct interaction with vulnerabilities after code changes, capabilities that existing large-scale datasets do not provide. In our evaluation, ARVO successfully reproduces 81% of vulnerabilities and achieves 89.4% accuracy on the located patches. We also discuss ARVO's influence on both upstream practices and downstream security research.
Xiang Mei, Jordi Del Castillo, Pulkit Singh Singaria, Haoran Xi, Abdelouahab Benchikh, Tiffany Bao, Ruoyu Wang 0001, Yan Shoshitaishvili, Adam Doupé, Hammond A. Pearce, Brendan Dolan-Gavitt
EuroS&P10
2026 Naming Variables is Hard - Assessing Names Need Not Be: An Inductive Taxonomy for Grading Variable Names
abstract
It is generally accepted amongst CS educators that creating ''good'' variable names is important. However, it is difficult to understand what makes a variable name ''good''. This makes it challenging to both educate and assess students on this topic. This is particularly true for high school teachers, many of whom are not experienced programmers. We therefore develop a decision-tree taxonomy of variable identifiers, with the goal of better equipping teachers to explain, understand, and assess what is ''good'' when it comes to their students' variable naming conventions and behaviours. Our taxonomy is based on 152 real high school student submissions to an automated assessment system used for a New Zealand high school programming standard in the Python programming language. From this, we perform an inductive analysis of 831 variable identifiers, and create a classification scheme made up of 25 distinct categories. We demonstrate that this taxonomy is usable outside of the developers, by scoring high inter-rater reliability with someone not involved in its development, and discuss how it could be adopted by high school educators.
Henry Hickman, Siyu Qiu, Hammond A. Pearce
ITiCSE (1)3
2025 What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
Jiamin Chang, Haoyang Li 0018, Hammond A. Pearce, Ruoxi Sun 0001, Bo Li 0026, Minhui Xue 0001
CCS3
2025 Logic Meets Magic: LLMs Cracking Smart Contract Vulnerabilities
ZeKe Xiao, Qin Wang 0008, Hammond A. Pearce, Shiping Chen 0001
ICBC3
2025 Experiences scaffolding a computer engineering project course to improve student outcomes
abstract
Computer engineering instruction at many academic institutions depends on open-ended discovery-based project courses. Depending upon the pre-existing knowledge of students, success rates can vary. This experiential paper describes how one such course, offered to third-year junior-level students, was updated to improve student outcomes by scaffolding the relevant knowledge and skills required to complete the digital design project. These updates resulted in both (1) significantly better student project outcomes (from 18% of groups completing the project to 83%), and (2) improved student satisfaction surveys for the course (rising 12.8 percentage points).
Hammond A. Pearce
ISCAS1
2025 Mitigation of Cyber-physical Attacks in Industry 4.0 using Secure Function Blocks
abstract
As Industry 4.0 drives the Fourth Industrial Revolution, Cyber-Physical Systems (CPSs) have become central to industrial automation. These systems integrate software with physical processes, significantly improving the efficiency and adaptability. However, this integration also expands the attack surface, exposing systems to Cyber-Physical-attacks (CP-attacks) that can target either the computational components, physical devices, or both. The impact of such attacks can be catastrophic, ranging from system disruption to physical damage. Although numerous techniques have been developed to detect and mitigate these threats, industrial standards are often not incorporated into the design of these methods. This limits their deployment within the Industry 4.0 systems, where standard compliance is critical.
Steph Wu, Nathan Allen, Alex Baird, Hammond A. Pearce, Partha S. Roop
MEMOCODE4
2025 High-Fidelity Specification of Real-World Devices
abstract
Device driver bugs are the leading cause of operating-system exploits, and the lack of accurate specifications of device interfaces is a leading cause of driver bugs. We propose to address the specification issue by deriving formal specifications of devices from their Verilog implementation, and prove the correctness of the specification against the implementation. We demonstrate this approach by applying it to an open-source I2C controller. These specifications should enable synthesis or verification of drivers in the future.
Liam Murphy 0003, Albert Rizaldi, Lesley Rossouw, Chen George, James Treloar, Hammond A. Pearce, Miki Tanaka, Gernot Heiser
PLOS@SOSP6
2025 FLAG: Finding Line Anomalies (in RTL code) with Generative AI
abstract
Bug detection in Hardware Design Languages (HDLs) is an important problem in the System-on-Chip (SoC) development cycle. It is crucial to find defects at the earliest stage possible. While most fault localization requires the use of “tests” (e.g., test benches, fuzzing, and assertions) and a simulation or emulation framework, the advent of Large Language Models (LLMs) provides an opportunity for a test-free fault localization approach. This article proposes such a tool, called FLAG, which can identify functional and security defects in Register Transfer Level (RTL) code without synthesis or simulation. FLAG combines syntactic and generative AI techniques to implement fault localization in RTL code. It takes an RTL design as an input and outputs a set of line(s) that likely contain defects. It targets elements of RTL code most likely to contain bugs through static analysis means and then implements token-level and line-level analysis to obtain differences in original code and code generated by LLM to identify a line as buggy or not. The token-level approach evaluates each generated token (one at a time) and the line level approach evaluates the entire line generated by the LLM. We evaluate our approach on a corpus of synthetic and real-world bugs, of both functional and security related issues, in Verilog and SystemVerilog. Using line-level analysis, FLAG can identify 38 out of 120 real-world bugs and using token-level analysis, FLAG can identify 32 out of 81 synthetic bugs through the top-5 most likely bug locations identified without tests.
Baleegh Ahmad, Joey Ah-kiow, Benjamin Tan 0001, Ramesh Karri, Hammond A. Pearce
ACM Trans. Design Autom. Electr. Syst.5
2025 Automatically Improving LLM-based Verilog Generation using EDA Tool Feedback
abstract
Traditionally, digital hardware designs are written in the Verilog hardware description language (HDL) and debugged manually by engineers. This can be time-consuming and error-prone for complex designs. Large Language Models (LLMs) are emerging as a potential tool to help generate fully functioning HDL code, but most works have focused on generation in the single-shot capacity: i.e., run and evaluate, a process that does not leverage debugging and, as such, does not adequately reflect a realistic development process. In this work, we evaluate the ability of LLMs to leverage feedback from electronic design automation (EDA) tools to fix mistakes in their own generated Verilog. To accomplish this, we present an open-source, highly customizable framework, AutoChip, which combines conversational LLMs with the output from Verilog compilers and simulations to iteratively generate and repair Verilog. To determine the success of these LLMs we leverage the VerilogEval benchmark set. We evaluate four state-of-the-art conversational LLMs, focusing on readily accessible commercial models. EDA tool feedback proved to be consistently more effective than zero-shot prompting only with GPT-4o, the most computationally complex model we evaluated. In the best case, we observed a 5.8% increase in the number of successful designs with a 34.2% decrease in cost over the best zero-shot results. Mixing smaller models with this larger model at the end of the feedback iterations resulted in equally as much success as with GPT-4o using feedback, but incurred 41.9% lower cost (corresponding to an overall decrease in cost over zero-shot by 89.6%).
Jason Blocklove, Shailja Thakur, Benjamin Tan 0001, Hammond A. Pearce, Siddharth Garg, Ramesh Karri
ACM Trans. Design Autom. Electr. Syst.4
2025 Introduction to Special Issue on Large Language Models for Electronic System Design Automation
abstract
Large Language Models are having a substantial impact on electronic design automation in areas ranging from hardware architecture to verification and optimization. The special issue provides a snapshot of work on this topic. This introduction describes and provides context for the research area, describes the organization of the special issue, and provides terse summaries of each of its papers.
Robert P. Dick, Hammond A. Pearce, Li Shang 0002, Fan Yang 0001
ACM Trans. Design Autom. Electr. Syst.2
2025 CGP-Tuning: Structure-Aware Soft Prompt Tuning for Code Vulnerability Detection
abstract
Large language models (LLMs) have been proposed as powerful tools for detecting software vulnerabilities, where task-specific fine-tuning is typically employed to provide vulnerability-specific knowledge to the LLMs. However, existing fine-tuning techniques often treat source code as plain text, losing the graph-based structural information inherent in code.Graph-enhanced soft prompt tuning addresses this by translating the structural information into contextual cues that the LLM can understand. However, current methods are primarily designed for general graph-related tasks and focus more on adjacency information, they fall short in preserving the rich semantic information (e.g., control/data flow) within code graphs. They also fail to ensure computational efficiency while capturing graph-text interactions in their cross-modal alignment module.This paper presents CGP-Tuning, a new code graph-enhanced, structure-aware soft prompt tuning method for vulnerability detection. CGP-Tuning introduces type-aware embeddings to capture the rich semantic information within code graphs, along with an efficient cross-modal alignment module that achieves linear computational costs while incorporating graph-text interactions. It is evaluated on the latestDiverseVuldataset and three advanced open-source code LLMs, CodeLlama, CodeGemma, and Qwen2.5-Coder. Experimental results show that CGP-Tuning delivers model-agnostic improvements and maintains practical inference speed, surpassing the best graph-enhanced soft prompt tuning baseline by an average of four percentage points and outperforming non-tuned zero-shot prompting by 15 percentage points.
Ruijun Feng, Hammond A. Pearce, Pietro Liguori, Yulei Sui
IEEE Trans. Software Eng.2
2024 Offramps: An FPGA-Based Intermediary for Analysis and Modification of Additive Manufacturing Control Systems
abstract
Cybersecurity threats in Additive Manufacturing (AM) are an increasing concern as AM adoption continues to grow. AM is now being used for parts in the aerospace, transportation, and medical domains. Threat vectors which allow for part compromise are particularly concerning, as any failure in these domains would have life-threatening consequences. A major challenge to investigation of AM part-compromises comes from the difficulty in evaluating and benchmarking both identified threat vectors as well as methods for detecting adversarial actions. In this work, we introduce a generalized platform for systematic analysis of attacks against and defenses for 3D printers. Our “OFFRAMPS” platform is based on the open-source 3D printer control board “RAMPS.“ Offramps allows analysis, recording, and modification of all control signals and I/O for a 3D printer. We show the efficacy of Offramps by presenting a series of case studies based on several Trojans, including ones identified in the literature, and show that Offramps can both emulate and detect these attacks, i.e., it can both change and detect arbitrary changes to the g-code print commands.
Jason Blocklove, Md Raz, Prithwish Basu Roy, Hammond A. Pearce, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri
DSN4
2024 dcc -help: Transforming the Role of the Compiler by Generating Context-Aware Error Explanations with Large Language Models
abstract
In the challenging field of introductory programming, high enrolments and failure rates drive us to explore tools and systems to enhance student outcomes, especially automated tools that scale to large cohorts. This paper presents and evaluates the dcc --help tool, an integration of a Large Language Model (LLM) into the Debugging C Compiler (DCC) to generate unique, novice-focused explanations tailored to each error. dcc --help prompts an LLM with contextual information of compile- and run-time error occurrences, including the source code, error location and standard compiler error message. The LLM is instructed to generate novice-focused, actionable error explanations and guidance, designed to help students understand and resolve problems without providing solutions. dcc --help was deployed to our CS1 and CS2 courses, with 2,565 students using the tool over 64,000 times in ten weeks. We analysed a subset of these error/explanation pairs to evaluate their properties, including conceptual correctness, relevancy, and overall quality. We found that the LLM-generated explanations were conceptually accurate in 90% of compile-time and 75% of run-time cases, but often disregarded the instruction not to provide solutions in code. Our findings, observations and reflections following deployment indicate that dcc --help provides novel opportunities for scaffolding students' introduction to programming.
Andrew Taylor, Alexandra Vassar, Jake Renzella, Hammond A. Pearce
SIGCSE (1)4
2024 LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
abstract
Large Language Models (LLMs) have been suggested for use in automated vulnerability repair, but benchmarks showing they can consistently identify security-related bugs are lacking. We thus develop SecLLMHolmes, a fully automated evaluation framework that performs the most detailed investigation to date on whether LLMs can reliably identify and reason about security-related bugs. We construct a set of 228 code scenarios and analyze eight of the most capable LLMs across eight different investigative dimensions using our framework. Our evaluation shows LLMs provide non-deterministic responses, incorrect and unfaithful reasoning, and perform poorly in real-world scenarios. Most importantly, our findings reveal significant non-robustness in even the most advanced models like ‘PaLM2’ and ‘GPT-4’: by merely changing function or variable names, or by the addition of library functions in the source code, these models can yield incorrect answers in 26% and 17% of cases, respectively. These findings demonstrate that further LLM advances are needed before LLMs can be used as general purpose security assistants.
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond A. Pearce, Ayse K. Coskun, Gianluca Stringhini
SP4
2024 REMaQE: Reverse Engineering Math Equations from Executables
abstract
Cybersecurity attacks on embedded devices for industrial control systems and cyber-physical systems may cause catastrophic physical damage as well as economic loss. This could be achieved by infecting device binaries with malware that modifies the physical characteristics of the system operation. Mitigating such attacks benefits from reverse engineering tools that recover sufficient semantic knowledge in terms of mathematical equations of the implemented algorithm. Conventional reverse engineering tools can decompile binaries to low-level code, but offer little semantic insight. This article proposes the REMaQE automated framework for reverse engineering of math equations from binary executables. Improving over state-of-the-art, REMaQE handles equation parameters accessed via registers, the stack, global memory, or pointers, and can reverse engineer equations from object-oriented implementations such as C++ classes. Using REMaQE, we discovered a bug in the Linux kernel thermal monitoring tool “tmon.” To evaluate REMaQE, we generate a dataset of 25,096 binaries with math equations implemented in C and Simulink. REMaQE successfully recovers a semantically matching equation for all 25,096 binaries. REMaQE executes in 0.48 seconds on average and in up to 2 seconds for complex equations. Real-time execution enables integration in an interactive math-oriented reverse engineering workflow.
Meet Udeshi, Prashanth Krishnamurthy, Hammond A. Pearce, Ramesh Karri, Farshad Khorrami
ACM Trans. Cyber Phys. Syst.3
2024 On Hardware Security Bug Code Fixes by Prompting Large Language Models
abstract
Novel AI-based code-writing Large Language Models (LLMs) such as OpenAI’s Codex have demonstrated capabilities in many coding-adjacent domains. In this work, we consider how LLMs may be leveraged to automatically repair identified security-relevant bugs present in hardware designs by generating replacement code. We focus on bug repair in code written in Verilog. For this study, we curate a corpus of domain-representative hardware security bugs. We then design and implement a framework to quantitatively evaluate the performance of any LLM tasked with fixing the specified bugs. The framework supports design space exploration of prompts (i.e., prompt engineering) and identifying the best parameters for the LLM. We show that an ensemble of LLMs can repair all fifteen of our benchmarks. This ensemble outperforms a state-of-the-art automated hardware bug repair tool on its own suite of bugs. These results show that LLMs have the ability to repair hardware security bugs and the framework is an important step towards the ultimate goal of an automated end-to-end bug repair tool.
Baleegh Ahmad, Shailja Thakur, Benjamin Tan 0001, Ramesh Karri, Hammond A. Pearce
IEEE Trans. Inf. Forensics Secur.5
2024 (Security) Assertions by Large Language Models
abstract
The security of computer systems typically relies on a hardware root of trust. As vulnerabilities in hardware can have severe implications on a system, there is a need for techniques to support security verification activities. Assertion-based verification is a popular verification technique that involves capturing design intent in a set of assertions that can be used in formal verification or testing-based checking. However, writing security-centric assertions is a challenging task. In this work, we investigate the use of emerging large language models (LLMs) for code generation in hardware assertion generation for security, where primarily natural language prompts, such as those one would see as code comments in assertion files, are used to produce SystemVerilog assertions. We focus our attention on a popular LLM and characterize its ability to write assertions out of the box, given varying levels of detail in the prompt. We design an evaluation framework that generates a variety of prompts, and we create a benchmark suite comprising real-world hardware designs and corresponding golden reference assertions that we want to generate with the LLM.
Rahul Kande, Hammond A. Pearce, Benjamin Tan 0001, Brendan Dolan-Gavitt, Shailja Thakur, Ramesh Karri, Jeyavijayan Rajendran
IEEE Trans. Inf. Forensics Secur.2
2024 VeriGen: A Large Language Model for Verilog Code Generation
abstract
In this study, we explore the capability of Large Language Models (LLMs) to automate hardware design by automatically completing partial Verilog code, a common language for designing and modeling digital systems. We fine-tune pre-existing LLMs on Verilog datasets compiled from GitHub and Verilog textbooks. We evaluate the functional correctness of the generated Verilog code using a specially designed test suite, featuring a custom problem set and testing benches. Here, our fine-tuned open-source CodeGen-16B model outperforms the commercial state-of-the-art GPT-3.5-turbo model with a 1.1% overall increase. Upon testing with a more diverse and complex problem set, we find that the fine-tuned model shows competitive performance against state-of-the-art gpt-3.5-turbo, excelling in certain scenarios. Notably, it demonstrates a 41% improvement in generating syntactically correct Verilog code across various problem categories compared to its pre-trained counterpart, highlighting the potential of smaller, in-house LLMs in hardware design automation. We release our training/evaluation scripts and LLM checkpoints as open-source contributions.
Shailja Thakur, Baleegh Ahmad, Hammond A. Pearce, Benjamin Tan 0001, Brendan Dolan-Gavitt, Ramesh Karri, Siddharth Garg
ACM Trans. Design Autom. Electr. Syst.3
2023 Benchmarking Large Language Models for Automated Verilog RTL Code Generation
abstract
Automating hardware design could obviate a signif-icant amount of human error from the engineering process and lead to fewer errors. Verilog is a popular hardware description language to model and design digital systems, thus generating Verilog code is a critical first step. Emerging large language models (LLMs) are able to write high-quality code in other programming languages. In this paper, we characterize the ability of LLMs to generate useful Verilog. For this, we fine-tune pre-trained LLMs on Verilog datasets collected from GitHub and Verilog textbooks. We construct an evaluation framework comprising test-benches for functional analysis and a flow to test the syntax of Verilog code generated in response to problems of varying difficulty. Our findings show that across our problem scenarios, the fine-tuning results in LLMs more capable of producing syntactically correct code (25.9% overall). Further, when analyzing functional correctness, a fine-tuned open-source CodeGen LLM can outperform the state-of-the-art commercial Codex LLM (6.5% overall). We release our training/evaluation scripts and LLM checkpoints as open source contributions.
Shailja Thakur, Baleegh Ahmad, Zhenxing Fan, Hammond A. Pearce, Benjamin Tan 0001, Ramesh Karri, Brendan Dolan-Gavitt, Siddharth Garg
DATE4
2023 Invited Paper: Towards the Imagenets of ML4EDA
abstract
Despite the growing interest in ML-guided EDA tools from RTL to GDSII, there are no standard datasets or prototypical learning tasks defined for the EDA problem domain. Experience from the computer vision community suggests that such datasets are crucial to spur further progress in ML for EDA. Here we describe our experience curating two large-scale, high-quality datasets for Verilog code generation and logic synthesis. The first, VeriGen, is a dataset of Verilog code collected from GitHub and Verilog textbooks. The second, OpenABC-D, is a large-scale, labeled dataset designed to aid ML for logic synthesis tasks. The dataset consists of 870,000 And-Inverter-Graphs (AIGs) produced from 1500 synthesis runs on a large number of open-source hardware projects. In this paper we will discuss challenges in curating, maintaining and growing the size and scale of these datasets. We will also touch upon questions of dataset quality and security, and the use of novel data augmentation tools that are tailored for the hardware domain.
Animesh Basak Chowdhury, Shailja Thakur, Hammond A. Pearce, Ramesh Karri, Siddharth Garg
ICCAD3
2023 An Integrated Testbed for Trojans in Printed Circuit Boards with Fuzzing Capabilities
abstract
This paper showcases an all-in-one testing environment that combines Trojan detection and fuzzing capabilities for printed circuit boards using the OpenPLC “NYU Trojan Edition” and a dedicated Trojan detection framework. The demo system is self-contained and equipped with two OpenPLC-based boards (one with a Trojan and one without), and automated tools for inserting the Trojan and collecting side-channel data. We developed a graphical user interface for interactive Trojan selection, data visualization, and anomaly detection analysis.
Prashanth Krishnamurthy, Hammond A. Pearce, Virinchi Roy Surabhi, Joshua Trujillo, Ramesh Karri, Farshad Khorrami
IOLTS2
2023 Examining Zero-Shot Vulnerability Repair with Large Language Models
abstract
Human developers can produce code with cybersecurity bugs. Can emerging ‘smart’ code completion tools help repair those bugs? In this work, we examine the use of large language models (LLMs) for code (such as OpenAI’s Codex and AI21’s Jurassic J-1) for zero-shot vulnerability repair. We investigate challenges in the design of prompts that coax LLMs into generating repaired versions of insecure code. This is difficult due to the numerous ways to phrase key information— both semantically and syntactically—with natural languages. We perform a large scale study of five commercially available, black-box, "off-the-shelf" LLMs, as well as an open-source model and our own locally-trained model, on a mix of synthetic, hand-crafted, and real-world security bug scenarios. Our experiments demonstrate that while the approach has promise (the LLMs could collectively repair 100% of our synthetically generated and hand-crafted scenarios), a qualitative evaluation of the model’s performance over a corpus of historical real-world examples highlights challenges in generating functionally correct code.
Hammond A. Pearce, Benjamin Tan 0001, Baleegh Ahmad, Ramesh Karri, Brendan Dolan-Gavitt
SP1
2023 Examining Zero-Shot Vulnerability Repair with Large Language Models
abstract
Human developers can produce code with cybersecurity bugs. Can emerging ‘smart’ code completion tools help repair those bugs? In this work, we examine the use of large language models (LLMs) for code (such as OpenAI’s Codex and AI21’s Jurassic J-1) for zero-shot vulnerability repair. We investigate challenges in the design of prompts that coax LLMs into generating repaired versions of insecure code. This is difficult due to the numerous ways to phrase key information— both semantically and syntactically—with natural languages. We perform a large scale study of five commercially available, black-box, "off-the-shelf" LLMs, as well as an open-source model and our own locally-trained model, on a mix of synthetic, hand-crafted, and real-world security bug scenarios. Our experiments demonstrate that while the approach has promise (the LLMs could collectively repair 100% of our synthetically generated and hand-crafted scenarios), a qualitative evaluation of the model’s performance over a corpus of historical real-world examples highlights challenges in generating functionally correct code.
Hammond A. Pearce, Benjamin Tan 0001, Baleegh Ahmad, Ramesh Karri, Brendan Dolan-Gavitt
SP1
2023 Lost at C: A User Study on the Security Implications of Large Language Model Code Assistants
Gustavo Sandoval, Hammond A. Pearce, Teo Nys, Ramesh Karri, Siddharth Garg, Brendan Dolan-Gavitt
USENIX Security Symposium2
2023 Multi-Modal Side Channel Data Driven Golden-Free Detection of Software and Firmware Trojans
abstract
This study explores data-driven detection of firmware/software Trojans in embedded systemswithoutgolden models. We consider embedded systems such as single board computers and industrial controllers. While prior literature considers side channel based anomaly detection, this study addresses the following central question: is anomaly detection feasible when using low-fidelity simulated data without using data from a known-good (golden) system? To study this question, we use data from a simulator-based proxy as a stand-in for unavailable golden data from a known-good system. Using data generated from the simulator, one-class classifier machine learning models are applied to detect discrepancies against expected side channel signal patterns and their inter-relationships. Side channels fused for Trojan detection include multi-modalside channelmeasurement data (such as Hardware Performance Counters, processor load, temperature, and power consumption). Additionally, fuzzing is introduced to increase detectability of Trojans. To experimentally evaluate the approach, we generate low-fidelity data using a simulator implemented with a component-based model and an information bottleneck based on Gaussian stochastic models. We consider example Trojans and show that fuzzing-aided golden-free Trojan detection is feasible using simulated data as a baseline.
Prashanth Krishnamurthy, Virinchi Roy Surabhi, Hammond A. Pearce, Ramesh Karri, Farshad Khorrami
IEEE Trans. Dependable Secur. Comput.3
2023 High-Level Approaches to Hardware Security: A Tutorial
abstract
Designers use third-party intellectual property (IP) cores and outsource various steps in the integrated circuit (IC) design and manufacturing flow. As a result, security vulnerabilities have been rising. This is forcing IC designers and end users to re-evaluate their trust in ICs. If attackers get hold of an unprotected IC, they can reverse engineer the IC and pirate the IP. Similarly, if attackers get hold of a design, they can insert malicious circuits or take advantage of “backdoors” in a design. Unintended design bugs can also result in security weaknesses. This tutorial paper provides an introduction to the domain of hardware security through two pedagogical examples of hardware security problems. The first is a walk-through of the scan chain-based side channel attack. The second is a walk-through of logic locking of digital designs. The tutorial material is accompanied by open access digital resources that are linked in this article.
Hammond A. Pearce, Ramesh Karri, Benjamin Tan 0001
ACM Trans. Embed. Comput. Syst.1
2022 Don't CWEAT It: Toward CWE Analysis Techniques in Early Stages of Hardware Design
abstract
To help prevent hardware security vulnerabilities from propagating to later design stages where fixes are costly, it is crucial to identify security concerns as early as possible, such as in RTL designs. In this work, we investigate the practical implications and feasibility of producing a set of security-specific scanners that operate on Verilog source files. The scanners indicate parts of code that might contain one of a set of MITRE's common weakness enumerations (CWEs). We explore the CWE database to characterize the scope and attributes of the CWEs and identify those that are amenable to static analysis. We prototype scanners and evaluate them on 11 open source designs - 4 system-on-chips (SoC) and 7 processor cores - and explore the nature of identified weaknesses. Our analysis reported 53 potential weaknesses in the OpenPiton SoC used in [email protected], 11 of which we confirmed as security concerns.
Baleegh Ahmad, Wei-Kai Liu, Luca Collini, Hammond A. Pearce, Jason M. Fung, Jonathan Valamehr, Mohammad Bidmeshki, Piotr Sapiecha, Krishnendu Chakrabarty, Ramesh Karri, Benjamin Tan 0001
ICCAD4
2022 Runtime Interchange of Enforcers for Adaptive Attacks: A Security Analysis Framework for Drones
abstract
Unmanned aerial drones are Cyber-Physical Systems (CPSs) with increasing availability, popularity, and capability. Although other aeronautical and safety-critical industries apply stringent regulations and design approaches, smaller drones tend to have much weaker and informal design requirements. Due to the strong open-source movement in this space, there are numerous opportunities for malicious actors to find weaknesses to attack drone systems, and in parallel develop their own rogue drones. These factors present a risk of damage to people and property in addition to compromise of integrity and availability. However, a formal framework for ethical hacking that combines attacker modelling and launching of attacks is lacking in the literature. To this end, we leverage runtime enforcement, combined with the idea of suspension from synchronous programming to develop the first such formal framework. The proposed framework enables the modelling of complex attack vectors on drones. To facilitate this, we propose a bespoke policy-based runtime enforcement framework called enforcer interchange (EI). It is capable of both individual intent/target-specific attacks as well as more sophisticated combinations of attacks, which it manages by enabling and disabling attack enforcers at runtime in a context-aware manner. To demonstrate our framework, we utilise a quadcopter drone simulator and record the changes in the drone's behaviour as it executes a range of missions under different attacks. Our approach provides a framework for testing drones' resilience and defenses against malicious attacks, as well as exploring the capabilities of rogue drones.
Alex Baird, Hammond A. Pearce, Srinivas Pinisetty, Partha S. Roop
MEMOCODE2
2022 Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
abstract
There is burgeoning interest in designing AI-based systems to assist humans in designing computing systems, including tools that automatically generate computer code. The most notable of these comes in the form of the first self-described ‘AI pair programmer’, GitHub Copilot, which is a language model trained over open-source GitHub code. However, code often contains bugs—and so, given the vast quantity of unvetted code that Copilot has processed, it is certain that the language model will have learned from exploitable, buggy code. This raises concerns on the security of Copilot’s code contributions. In this work, we systematically investigate the prevalence and conditions that can cause GitHub Copilot to recommend insecure code. To perform this analysis we prompt Copilot to generate code in scenarios relevant to high-risk cybersecurity weaknesses, e.g. those from MITRE’s “Top 25” Common Weakness Enumeration (CWE) list. We explore Copilot’s performance on three distinct code generation axes—examining how it performs given diversity of weaknesses, diversity of prompts, and diversity of domains. In total, we produce 89 different scenarios for Copilot to complete, producing 1,689 programs. Of these, we found approximately 40% to be vulnerable.
Hammond A. Pearce, Baleegh Ahmad, Benjamin Tan 0001, Brendan Dolan-Gavitt, Ramesh Karri
SP1
2022 Detecting Hardware Trojans in PCBs Using Side Channel Loopbacks
abstract
Malicious modifications to printed circuit boards (PCBs) are known as hardware Trojans. These may arise when malafide third parties alter PCBs premanufacturing or postmanufacturing and are a concern in safety-critical applications, such as industrial control systems. In this research, we examine how data-driven detection can be utilized to detect such Trojans at run-time. We develop a flexible and reconfigurable PCB test bed derived from the popular open-source programmable logic controller (PLC) platform “OpenPLC.” We then develop a Trojan detection framework, which utilizes and analyzes multimodal side channels (e.g., timing, magnetic signals, power, and hardware performance counters). We consider defender-configurable input/output (I/O) loopback test, comparison with design-document baselines, and magnetometer-aided monitoring of system behavior under defender-chosen excitations. Our approach can extend to golden-free environments. Golden (known-good) versions of the PCBs are assumed not available, but design information, datasheets, and component-level data are available. We demonstrate the efficacy of our approach on a range of Trojans instantiated in the test bed.
Hammond A. Pearce, Virinchi Roy Surabhi, Prashanth Krishnamurthy, Joshua Trujillo, Ramesh Karri, Farshad Khorrami
IEEE Trans. Very Large Scale Integr. Syst.1
2020 A compositional approach using Keras for neural networks in real-time systems
abstract
Real-time systems are designed using model-driven approaches, where a complex system is represented as a set of interacting components. Such a compositional approach facilitates design of simpler components, which are easier to validate and integrate with the overall system. In contrast to such systems, data-driven systems like neural networks are designed as monolithic black-boxes to capture the non-linear relationship from inputs to outputs. Increasingly, such systems are being used in safety-critical real-time systems. Here, a compositional approach would be ideal. However, to the best of our knowledge, such a compositional approach is lacking while designing data-driven components based on neural networks.This paper formalises this problem by developing the concept of Composed Neural Networks (CpNNs) by extending the well known Keras python framework. CpNNs formalise the synchronous composition of several interacting neural networks in Keras. Further, using the developed semantics, we enable modular compilation from a given CpNN to C code. The generated code is suitable for the Worst-Case Execution Time (WCET) analysis. Using several benchmarks we demonstrate the superiority of the developed approach over a recently proposed approach using Esterel, as well as the popular Python package Tensorflow Lite. For the given benchmarks, our approach is superior to Esterel with an average WCET reduction of 64.06%, and superior to Tensorflow Lite with an average measured WCET reduction of 62.08%.
Xin Yang 0035, Partha S. Roop, Hammond A. Pearce, Jin Woo Ro
DATE3
2020 Smart I/O Modules for Mitigating Cyber-Physical Attacks on Industrial Control Systems
abstract
Cyber-physical systems (CPSs) are implemented in many industrial and embedded control applications. Where these systems are safety-critical, correct and safe behavior is of paramount importance. Malicious attacks on such CPSs can have far-reaching repercussions. For instance, if elements of a power grid behave erratically, physical damage and loss of life could occur. Currently, there is a trend toward increased complexity and connectivity of CPS. However, as this occurs, the potential attack vectors for these systems grow in number, increasing the risk that a given controller might become compromised. In this article, we examine how the dangers of compromised controllers can be mitigated. We propose a novel application of runtime enforcement that can secure the safety of real-world physical systems. Here, we synthesize enforcers to a new hardware architecture within programmable logic controller I/O modules to act as an effective line of defence between the cyber and the physical domains. Our enforcers prevent the physical damage that a compromised control system might be able to perform. To demonstrate the efficacy of our approach, we present several benchmarks, and show that the overhead for each system is extremely minimal.
Hammond A. Pearce, Srinivas Pinisetty, Partha S. Roop, Matthew M. Y. Kuo, Abhisek Ukil
IEEE Trans. Ind. Informatics1
2019 Securing implantable medical devices with runtime enforcement hardware
abstract
In recent years we have seen numerous proof-of-concept attacks on implantable medical devices such as pacemakers. Attackers aim to breach the strict operational constraints that these devices operate within, with the end-goal of compromising patient safety and health. Most efforts to prevent these kinds of attacks are informal, and focus on application- and system-level security --- for instance, using encrypted communications and digital certificates for program verification. However, these approaches will struggle to prevent all classes of attacks. Runtime verification has been proposed as a formal methodology for monitoring the status of implantable medical devices. Here, if an attack is detected a warning is generated. This leaves open the risk that the attack can succeed before intervention can occur. In this paper, we propose a runtime-enforcement based approach for ensuring patient security. Custom hardware is constructed for individual patients to ensure a safe minimum quality of service at all times. To ensure correctness we formally verify the hardware using a model-checker. We present our approach through a pacemaker case study and demonstrate that it incurs minimal overhead in terms of execution time and power consumption.
Hammond A. Pearce, Matthew M. Y. Kuo, Partha S. Roop, Srinivas Pinisetty
MEMOCODE1
2018 Faster Function Blocks for Precision Timed Industrial Automation
abstract
In industrial automation, safety-critical control systems need robust timing guarantees in addition to functional correctness. Unfortunately, devices that are typically used in this domain, such as Programmable Logic Controllers, often feature architectures that are not amenable to static timing analysis, for instance relying on general purpose microprocessors or embedded operating systems. As a result, designers often rely on timing values gained from simple measurement of running applications, an approach that only provides very weak guarantees at best. The synchronous approach for IEC 61499 Function Blocks, in contrast, has been demonstrated to be time predictable when run on appropriate hardware, such as simple microprocessors. However, simple microprocessors are often not fast or powerful enough for modern automation requirements. In this paper, we examine how the performance of synchronous IEC 61499 can be improved through the usage of the multi-core T-CREST architecture, data scratchpads, and an optimised compiler. Overall, our improvements resulted in 60% shorter worst-case execution times.
Hammond A. Pearce, Partha S. Roop, Morteza Biglari-Abhari, Martin Schoeberl
ISORC1
2018 Synchronous neural networks for cyber-physical systems
abstract
Cyber-physical systems (CPS), such as autonomous vehicles or smart power grids, use interactive machine learning modules for decision making. Current design approaches use multiple machine learning modules, often using neural networks, to achieve the desired functionality. Timing validation is performed using measurement-based approaches, which may produce unsound results. To this end, we propose a new approach for designing such systems, by relying on the well known synchronous paradigm. Using this approach, we introduce Synchronous Artificial Neural Networks (SANNs), where we associate logical time to the different operations of the network. This approach provides sound compositional primitives, which enable the composition of interacting neural networks to ensure causality and determinism. We then show that we can embed the generated code on time predictable platforms enabling static analysis. Overall, this paper develops synchronous neural networks implemented in Esterel for the design of time predictable systems. We demonstrate the efficacy of our approach by developing a time predictable implementation of several applications, ranging from 5-100+ neurons, realised using the T-CREST platform. We also implemented a complex Convolutional Neural Network (CNN) application comprising of 1000+ neurons and 16 different layers using Esterel for a soft real-time application. Overall, the developed methodology opens new avenues of research in the direction of time predictable neural networks.
Partha S. Roop, Hammond A. Pearce, Keyan Monadjem
MEMOCODE2
2017 A Model Driven Approach for Cardiac Pacemaker Design Using a PRET Processor
abstract
Implantable medical devices such as cardiac pacemakers have been recalled frequently with safety related issues. This paper proposes a model driven approach for pacemaker design by combining the strengths of two well-known philosophies for safety critical systems. First, we adopt the SCCharts synchronous language for pacemaker specification. Second, we adopt a PRET architecture for the underlying processor which has been modified to include reactive semantics. PRET processors offer an ideal platform for providing timing guarantees. We use automatic code generation combined with static timing analysis during the design phase. Also, we use an existing emulation model of the human heart using a 33-node conduction network for closed loop validation of the designed pacemaker.
Nathan Allen, Hammond A. Pearce, Partha S. Roop, Reinhard von Hanxleden
ISORC2
2017 Simulation of cyber-physical systems using IEC61499
abstract
IEC61499 is an emerging standard for the design of automation systems. While many compilers and associated tools for IEC61499 have been developed, systematic techniques for modelling the continuous dynamics of the physical processes are lacking. Current practices involve using co-simulation, where plants are modelled in a tool such as Simulink and controllers are designed using IEC-61499. Co-simulation has many limitations such as slow sampling and free-wheeling. In this paper we propose a systematic approach for the design and simulation of Cyber-Physical Systems (CPS) using IEC61499. We propose the concept of Hybrid Function Blocks (HFBs), as syntactic extensions, to specify the continuous dynamics of a physical plant. A Hybrid Function Block can be compiled into a standards compliant Basic Function Block, based on new deterministic synchronous semantics. To show that our approach is both scalable and efficient when designing CPS, we present benchmarks showing that it runs 29 % faster than Simulink when generating correlating traces.
Hammond A. Pearce, Matthew M. Y. Kuo, Nathan Allen, Partha S. Roop, Avinash Malik
MEMOCODE1
2016 RunSync: A Predictable Runtime for Precision Timed Automation Systems
abstract
Many complex industrial control systems need to meet stringent timing requirements. Ensuring that implementations can meet these requirements can be a very difficult task. IEC 61499 is an emerging standard for modelling and implementing large distributed industrial control systems. However, there is currently no established approach for executing IEC 61499 code in a distributed time-predictable manner. In this paper, we present a novel time-predictable runtime called RunSync for executing IEC 61499 on the Precision Timed (PRET) architecture FlexPRET. PRET architectures are designed to guarantee timing repeatability while preserving performance. They are often designed as RISC architectures which utilize multiple hardware threads to remove pipeline hazards. Determining the allocation of tasks to the hardware threads is a key problem when utilizing such architectures. Hence, in this paper it is demonstrated that through RunSync it is possible to dynamically map IEC 61499 tasks to hardware threads during runtime, while preserving determinism and remaining amenable to timing analysis. Following that, quantitative results are presented, showing the minimal overheads of RunSync compared to implementations of the existing synchronous approach. RunSync is thus demonstrated to be more performant with large IEC 61499 networks, and when there are more hardware resources to be allocated.
Hammond A. Pearce, Matthew M. Y. Kuo, Partha S. Roop, Morteza Biglari-Abhari
ISORC1