Soheil Salehi

dblp:166/3227 · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0001-5998-8795ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 6 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight
abstract
Emotional coordination is a core property of human interaction that shapes how relational meaning is constructed in real time. While text-based affect inference has become increasingly feasible, prior approaches often treat sentiment as a deterministic point estimate for individual speakers, failing to capture the inherent subjectivity, latent ambiguity, and sequential coupling found in mutual exchanges. We introduce LLM-MC-Affect, a probabilistic framework that characterizes emotion not as a static label, but as a continuous latent probability distribution defined over an affective space. By leveraging stochastic LLM decoding and Monte Carlo estimation, the methodology approximates these distributions to derive high-fidelity sentiment trajectories that explicitly quantify both central affective tendencies and perceptual ambiguity. These trajectories enable a structured analysis of interpersonal coupling through sequential cross-correlation and slope-based indicators, identifying leading or lagging influences between interlocutors. To validate the interpretive capacity of this approach, we utilize teacher-student instructional dialogues as a representative case study, where our quantitative indicators successfully distill high-level interaction insights such as effective scaffolding. This work establishes a scalable and deployable pathway for understanding interpersonal dynamics, offering a generalizable solution that extends beyond education to broader social and behavioral research.
Yu-Zheng Lin, Bono Po-Jen Shih, John Paul Martin Encinas, Elizabeth Victoria Abraham Achom, Karan Himanshu Patel, Jesus Horacio Pacheco, Sicong Shao, Jyotikrishna Dass, Soheil Salehi, Pratik Satam
ACL (1)9
2026 Photogrammetry-Enabled Digital Twins for Semiconductor Education and Workforce Development
abstract
Semiconductor manufacturing is a critical and vital industry impacting all aspects of modern life centered around artificial intelligence (AI). The demand for semiconductors has skyrocketed with the emergence of the Fourth Industrial Revolution (4IR) technologies that integrate traditional manufacturing with modern technologies such as cloud computing, machine learning, and artificial intelligence. These transformations have created an ever-increasing need to manufacture semiconductors, creating a massive investment and the need for new fabrication capabilities over the next decade, presenting an enormous workforce development challenge. The complexity of the semiconductor manufacturing environment and equipment used exacerbates this challenge, making this workforce development a pressing need. Although not new, photogrammetry has proven transformative in industry and education, providing customizable visual digital twins of processes, facilities, and equipment. This paper highlights our ongoing efforts to use photogrammetry to build 3D models and level 1 Digital Twins, that will integrate into our PRISM platform. To facilitate this curriculum development effort, we aim to build experiential learning exercises in the PRISM platform, focusing on equipment setup, equipment usage, and basic debugging and problem solving, capitalizing on the PRISM platform’s ability to measure student sentiment while performing experiential learning to personalize and generate new content for their training needs using Retrieval-Augmented Generation (RAG) and generative AI. Our goal is to model different different stages of semiconductor manufacturing, namely 1) Photoresist Application, 2) Lithography, 3) Nickel Chromium Deposition, and 4) Lift-off. To achieve this goal, we are modeling Laurell Spin Coater, ABM Mask Aligner, Thermal Physical Vapor Evaporator, Ellipsometer, and a Profilometer using Photogrammetry. On integration into the PRISM platform, capitalizing on the PRISM platform’s ability to personalize training content through generative AI via student sentiment analysis for different target student groups, we aim to rapidly create semiconductor training experiential learning exercises for High School, Undergraduate, and Graduate Students, helping create a specialized semiconductor workforce to meet the current pressing semiconductor workforce needs.
John Paul Martin Encinas, Yu-Zheng Lin, Bono Po-Jen Shih, Josh Dean, Anh Minh Nguyen, Aurora Namjoshi, Ahmed Alhamadah, Shalaka Satam, Soheil Salehi, Pratik Satam
ACM Great Lakes Symposium on VLSI10
2026 LeakSEAL: Power Side-Channel Leakage Analysis and Mitigation for Secure Edge AI Learning
abstract
On-chip learning enables machine learning models to be trained or updated directly on specialized hardware rather than on external CPUs or GPUs, offering lower latency, improved energy-efficiency, enhanced privacy, and real-time adaptability for edge devices. In Spiking Neural Networks (SNNs), this capability relies on dynamic synaptic weight adaptation, but such adaptability also introduces significant security risks. In this work, we demonstrate a power side-channel attack on a quantized SNN implemented on a CW305 FPGA platform using ChipWhisperer. Our analysis identifies consistent power leakage patterns associated with neuron update operations, allowing an attacker to infer internal model attributes without direct access to the model’s weights or inputs. We further perform Correlation Power Analysis (CPA) with a Hamming Weight leakage model to recover secret synaptic weights with high confidence using as few as 1,500 power traces. These results expose critical vulnerabilities in on-chip learning systems and SNN architectures, highlight realistic threats to IoT and edge applications, and motivate mitigation strategies at the software-hardware boundary, including secure design practices, cryptographic protections, and access control mechanisms, without significantly degrading performance.
Veeramani Pugazhenthi, Md Muhtasim Alam Chowdhury, Sujan Ghimire, Harish Kumar Dharavath, Parsa Mirfasihi, Nader Sehatbakhsh, Pratik Satam, Soheil Salehi
ACM Great Lakes Symposium on VLSI8
2026 Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection
Sujan Ghimire, Parsa Mirfasihi, Md Muhtasim Alam Chowdhury, Veeramani Pugazhenthi, Harish Kumar Dharavath, Farshad Firouzi, Rozhin Yasaei, Pratik Satam, Soheil Salehi
VTS9
2025 Transformers for Secure Hardware Systems: Applications, Challenges, and Outlook
Banafsheh S. Latibari, Najmeh Nazari, Avesta Sasan, Houman Homayoun, Pratik Satam, Soheil Salehi, Hossein Sayadi
ACM Great Lakes Symposium on VLSI6
2025 Llm4mcu-Onto: Leveraging Llms for Automated Ontology Generation From Microcontroller Reference Manual
abstract
This research addresses the challenges faced by firmware developers, security researchers, and enthusiasts who work with low-level microcontroller (MCU) documentation, which often spans hundreds of complex pages. Current structured approaches, such as System View Description (SVD), are widely used but suffer from manual, labor-intensive creation processes and inconsistent vendor adherence to CMSIS-SVD standards. We propose an automated solution using Large Language Models (LLMs) integrated with Retrieval-Augmented Generation (RAG), capable of effectively parsing and extracting structured information, including text, tables, and images from MCU reference manuals/ datasheets. To mitigate hallucination issues inherent in LLMs, we fine-tuned models using a dataset derived from CMSIS-SVD files, which we will open-source for community benefit. We also experimented with few-shot models. Additionally, we developed a standardized structured ontology that is automatically populated with information extracted through LLM assistance from the reference manuals of the corresponding MCUs. Our approach was evaluated using OpenAI's GPT-4o under one-shot, few-shot, and fine-tuning scenarios, all incorporating RAG. We also experimented with the open-source LLM model CodeLlama. The results highlight substantial improvements in automatically extracting peripheral details and information from MCU reference manuals. Thus, it helps reduce manual effort and time. The key contribution of our work lies in the tailored adaptation of existing AI techniques to address the specific challenges of embedded systems documentation. We perform standardized ontology creation and multimodal parsing. We leverage RAG with MCU-specific finetuning and few-shot learning to generate structured information from hundreds of pages of MCU documentation. This opens the door to potential applications such as more accurate firmware code generation and reverse engineering for security analysis.
Asmita 0001, Grisha Bandodkar, Sujan Ghimire, Shaurya Srivastav, Soheil Salehi, Houman Homayoun
ICCD5
2025 Reliability of Capacitive Read in Arrays of Ferroelectric Capacitors
abstract
The non-destructive capacitance read-out of ferroelectric capacitors (FeCaps) based on doped HfO2metal-ferroelectric-metal (MFM) structures offers the potential for low-power and highly scalable crossbar arrays. This is due to a number of factors, including the selector-less design, the absence of sneak paths, the power-efficient charge-based read operation, and the reduced IR drop. Nevertheless, a reliable capacitive readout presents certain challenges, particularly in regard to device variability and the trade-off between read yield and read disturbances, which can ultimately result in bit-flips. This paper presents a digital read macro for HfO2FeCaps and provides reliability analysis for the capacitive readout of HfO2FeCaps, taking device variability and yield challenges into account. An experimentally calibrated physics-based compact model of HfO2FeCaps is employed to investigate the reliability of the read-out operation of the FeCap macro through Monte Carlo simulations. Based on this analysis, we identify limitations posed by the device variability and propose potential mitigation strategies through design-technology co-optimization (DTCO) of the FeCap device characteristics and the CMOS circuit design. Finally, we examine the potential applications of the FeCap macro in the context of secure hardware. We identify potential security threats and propose strategies to enhance the robustness of the system.
Luca Fehlings, Md Muhtasim Alam Chowdhury, Banafsheh S. Latibari, Soheil Salehi, Erika Covi
ISCAS4
2025 HWREx: AI-enabled Hardware Weakness and Risk Exploration and Storytelling Framework with LLM-assisted Mitigation Suggestion
abstract
The growing complexity of modern computing frameworks has led to an increase in cybersecurity vulnerabilities reported to the National Vulnerability Database (NVD). Extracting meaningful trends from this vast amount of unstructured data is challenging without proper tools and methodologies. Existing approaches lack a holistic strategy for vulnerability mitigation and prediction and effective knowledge extraction from the Common Weakness Enumeration (CWE), Common Vulnerability Exposure (CVE), and Common Attack Pattern Enumeration and Classification (CAPEC) databases. We introduce the AI-enabled Hardware Weakness and Risk Exploration and Storytelling Framework with LLM-assisted Mitigation Suggestion (HWREx), designed to address hardware vulnerabilities and IoT security. Our architecture features an Ontology-driven Storytelling capability that automates ontology updates to track vulnerability patterns and evolution over time, while offering mitigation strategies. It also clarifies the complex interrelations among CVEs, CWEs, and CAPECs through interactive visual knowledge graphs. Our framework achieved accuracy rates of 62% for CWE-CWE, 83% for CWE-CVE, and 77% for CWE-CAPEC linkage predictions. These graphs are instrumental for in-depth hardware weakness analysis and enable HWREx to deliver comprehensive assessments and actionable mitigation strategies. Additionally, HWREx utilizes Generative Pre-trained Transformers (GPT) to offer tailored mitigation suggestions.
Sujan Ghimire, Yu-Zheng Lin, Muntasir Mamun, Md Muhtasim Alam Chowdhury, Farhad Alemi, Shuyu Cai, Jinduo Guo, Banafsheh S. Latibari, Setareh Rafatirad, Pratik Satam, Soheil Salehi
ACM Trans. Design Autom. Electr. Syst.13
2024 Photogrammetry for Digital Twinning Industry 4.0 (I4) Systems
abstract
The onset of Industry 4.0 is rapidly transforming the manufacturing world through the integration of cloud computing, machine learning (ML), artificial intelligence (AI), and universal network connectivity, resulting in performance optimization and increased productivity. Digital Twins (DT) are one such transformational technology that leverages software systems to replicate physical process behavior, and representing it in a digital environment. This paper aims to explore the use of photogrammetry (which is the process of reconstructing physical objects into virtual 3D models using photographs) and 3D Scanning techniques to create accurate visual representation of the ‘Physical Process', to interact with the ML/AI based behavior models. To achieve this, we have used a readily available consumer device, the iPhone 15 Pro, which features stereo vision capabilities, to capture the depth of an Industry 4.0 system. By processing these images using 3D scanning tools, we created a raw 3D model for 3D modeling and rendering software for the creation of a DT model. The paper highlights the reliability of this method by measuring the error rate in between the ground truth (measurements done manually using a tape measure) and the final 3D model created using this method. The overall mean error is 4.97 % and the overall standard deviation error is 5.54% between the ground truth measurements and their photogrammetry counterparts. The results from this work indicate that photogrammetry using consumer-grade devices can be an efficient and cost-efficient approach to creating DTs for smart manufacturing, while the approaches flexibility allows for iterative improvements of the models over time.
Ahmed Alhamadah, Muntasir Mamun, Henry Harms, Mathew Redondo, Yu-Zheng Lin, Soheil Salehi, Pratik Satam
AICCSA7
2024 Interactive Framework for Cybersecurity Education and Future Workforce Development
abstract
This research-to-practice paper presents a novel pedagogical tool for hardware cybersecurity education and workforce development. The growing importance of hardware security has made it essential for individuals and organizations to understand hardware security principles and best practices. However, the current educational curriculum falls short of fulfilling these emerging demands due to the rapidly changing hardware security landscape and limited opportunities for hands-on training. To address these challenges, we propose and have developed the Interactive Hardware and Cybersecurity (I-HaC) Educational Framework, a pedagogical educational framework that supplements existing courses by leveraging generative AI for individualized instruction related to hardware and cybersecurity, data mining, and applied Machine Learning (ML), as well as data visualization to enhance cybersecurity education and workforce development. The framework is designed to be utilized by graduate and undergraduate Electrical and Computer Engineering (ECE) and Computer Science (CS) students for a comprehensive introduction to cybersecurity exploits and countermeasures in an interactive manner with hands-on components. Using I-HaC, we have developed tailored lab components for a diverse range of students and intend to release I-HaC as open-source for the benefit of the ECE and CS education community.
Sujan Ghimire, Md Muhtasim Alam Chowdhury, Ryan Tsang, Richard C. Yarnell, Emma Heckert, Jaeden Wolf Carpenter, Yu-Zheng Lin, Muntasir Mamun, Ronald F. DeMara, Setareh Rafatirad, Pratik Satam, Soheil Salehi
FIE12
2024 IRET: Incremental Resolution Enhancing Transformer
abstract
In our research paper, we introduce a revolutionary approach to designing energy-aware dynamically prunable Vision Transformers for use in edge applications. Our solution denoted as Incremental Resolution Enhancing Transformer (IRET), works by the sequential sampling of the input image. However, in our case, the embedding size of input tokens is considerably smaller than prior-art solutions. This embedding is used in the first few layers of the IRET vision transformer until a reliable attention matrix is formed. Then the attention matrix is used to sample additional information using a learnable 2D lifting scheme only for important tokens and IRET drops the tokens receiving low attention scores. Hence, as the model pays more attention to a subset of tokens for its task, its focus and resolution also increase. This incremental attention-guided sampling of input and dropping of unattended tokens allow IRET to significantly prune its computation tree on demand. By controlling the threshold for dropping unattended tokens and increasing the focus of attended ones, we can train a model that dynamically trades off complexity for accuracy. This is especially useful for edge devices, where accuracy and complexity could be dynamically traded based on factors such as battery life, reliability, etc.
Banafsheh S. Latibari, Soheil Salehi, Houman Homayoun, Avesta Sasan
ACM Great Lakes Symposium on VLSI2
2024 Educational Tool-spaces for Convolutional Neural Network FPGA Design Space Exploration Using High-Level Synthesis
abstract
There is significant demand and urgency to prepare electrical and computer engineering students regarding the operational and performance characteristics of machine learning (ML) hardware accelerators. Convolutional Neural Networks (CNNs), which are utilized for real-time and large dataset image classification tasks, are appropriate targets for hardware acceleration. Designing accelerators for CNNs necessitates understanding the manipulation of CNN parameters. We introduce a hands-on pedagogy whereby learners can identify, modify, and appreciate the interaction of the CNN parameters within an interactive GUI. CASCADE (Computer Aided Student's CNN Analyzer for Design Exploration), a simulation-based framework for Design Space Exploration (DSE) of CNN FPGA-based accelerators is developed, including datapath synthesis, simulation, training, and testbench steps. We offer a case study of High-Level Synthesis (HLS) based CNN implementations targeting the MNIST dataset and present simulation results, namely hardware utilization, accuracy, and operating frequency, and offer insight into potential design trade-offs facing modern engineers.
Richard C. Yarnell, Mousam Hossain, Raul Graterol, Ayush Pindoria, Sujan Ghimire, Md Muhtasim Alam Chowdhury, Soheil Salehi, Yu Bai 0004, Ronald F. DeMara
ACM Great Lakes Symposium on VLSI7
2024 Securing On-Chip Learning: Navigating Vulnerabilities and Potential Safeguards in Spiking Neural Network Architectures
abstract
On-chip learning is the process of training or updating machine learning models directly on specialized hardware. This approach differs from traditional machine learning, which typically conducts training on external computing resources like Central Processing Units (CPUs) or Graphics Processing Units (GPUs). On-chip learning offers several advantages, including reduced latency, improved energy efficiency, enhanced privacy, and adaptability. Consequently, it holds great promise for enabling intelligent decision-making and adaptability in resource-constrained edge and IoT devices while addressing privacy concerns. In Spiking Neural Network (SNN), on-chip learning is enabled by adjusting synaptic weights, allowing the network’s behavior to dynamically align with desired outcomes. However, this adaptability may introduce potential security vulnerabilities. Unmitigated security risks in on-chip learning can lead to various threats, including data leaks, unauthorized access, and even adversarial manipulation of the learning process. This manuscript aims to provide a comprehensive overview of the security risks associated with on-chip learning, with a focus on potential vulnerabilities within the SNN architecture. We will explore real-world scenarios where these vulnerabilities can be exploited and outline protective measures and mitigation strategies to address these security concerns.
Najmeh Nazari, Kevin Immanuel Gubbi, Banafsheh S. Latibari, Md Muhtasim Alam Chowdhury, Chongzhou Fang, Avesta Sasan, Setareh Rafatirad, Houman Homayoun, Soheil Salehi
ISCAS9
2024 FFXE: Dynamic Control Flow Graph Recovery for Embedded Firmware Binaries
Ryan Tsang, Asmita 0001, Doreen Joseph, Soheil Salehi, Prasant Mohapatra, Houman Homayoun
USENIX Security Symposium4
2024 Optimized and Automated Secure IC Design Flow: A Defense-in-Depth Approach
abstract
The globalization of the manufacturing process and the supply chain for electronic hardware has been driven by the need to maximize profitability while lowering risk in a technologically advanced silicon sector. However, many hardware IPs’ security features have been broken because of the rise in successful hardware attacks. Existing security efforts frequently ignore numerous dangers in favor of fixing a particular vulnerability. This inspired the development of a unique method that uses emerging spin-based devices to obfuscate circuitry to secure hardware intellectual property (IP) during fabrication and the supply chain. We propose an Optimized and Automated Secure IC (OASIC) Design Flow, a defense-in-depth approach that can minimize overhead while maximizing security. Our EDA tool flow uses a dynamic obfuscation method that employs dynamic lockboxes, which include switch boxes and magnetic random access memory (MRAM)-based look-up tables (LUT) while offering minimal overhead and being flexible and resilient against modern SAT-based attacks and power side-channel attacks. An EDA tool flow for optimized lockbox insertion is also developed to generate SAT-resilient design netlists with the least power and area overhead. PPA metrics and security (SAT attack time) are provided to the designer for each lockbox insertion run. A verification methodology is provided to verify locked and unlocked designs for functional correctness. Finally, we use ISCAS’85 benchmarks to show that the EDA tool flow provides a secure hardware netlist with maximum security while considering power and area constraints. Our results indicate that the proposed OASIC design flow can maximize security while incurring less than 15% area overhead and maintaining a similar power footprint compared to the original design. OASIC design flow demonstrates improved performance as design size increases, which demonstrates the scalability of the proposed approach.
Kevin Immanuel Gubbi, Banafsheh S. Latibari, Md Muhtasim Alam Chowdhury, Afrooz Jalilzadeh, Erfan Yazdandoost Hamedani, Setareh Rafatirad, Avesta Sasan, Houman Homayoun, Soheil Salehi
IEEE Trans. Circuits Syst. I Regul. Pap.9
2023 Machine Learning for Intrusion Detection: Stream Classification Guided by Clustering for Sustainable Security in IoT
abstract
The Internet of Things (IoT) has brought about unprecedented connectivity and convenience in our daily lives, but with this newfound interconnectedness comes the threat of cyber-attacks. With ever-increasing IoT devices being connected to the internet, securing IoT devices is becoming increasingly urgent. Machine learning (ML) is among the most popular techniques used by intrusion detection systems (IDS) to enhance their detection performance when securing IoT. However, a key obstacle of ML-based IDS for IoT is learning from nonstationary streaming data, also known as concept drift. One of the most challenging learning scenarios under concept drift is extreme verification latency (EVL), which occurs when only unlabeled nonstationary streaming data is available after a small set of initial labeled data. Stream Classification Algorithm Guided by Clustering (SCARGC) is an algorithm that can effectively deal with the nonstationary data streams in EVL scenarios. Applying an EVL implementation provides the capability of adapting to nonstationary environments within the IoT domain. The SCARGC model, as an integrated IoT intrusion detection system, allows for sustainable security as new threats are identified in this non-stationary environment. Hence, in this project, we develop an innovative IoT intrusion detection approach by natively integrating SCARGC and intrusion detection to address the EVL challenges to provide sustainable security as the model adapts to nonstationary environments. We evaluated the proposed approach on real-world IoT cybersecurity datasets. The results demonstrate the feasibility of the proposed approach, which can lead to the development of sophisticated intrusion detection systems for IoT.
Martin Manuel Lopez, Sicong Shao, Salim Hariri, Soheil Salehi
ACM Great Lakes Symposium on VLSI4
2023 Leveraging Firmware Reverse Engineering for Stealthy Sensor Attacks via Binary Modification
abstract
The number of Internet of Things (IoT) devices has increased dramatically to the point where they pervade our daily life. These connected devices are equipped with a variety of sensors for applications ranging from simple thermostats to critical medical devices. These devices often directly interact with people and usually lack proper security measures, thus they have become ideal targets for attackers. Herein, we propose Cunning Sensor Attack via Firmware Reverse-Engineering (unSAFE), which is a novel and stealthy sensor attack that attempts to corrupt sensor data by targeting the device’s Power Management IC (PMIC) configuration in firmware. The proposed unSAFE explores a class of vulnerabilities in which firmware is used to launch a physical attack against a device’s peripherals utilizing power management units as a vector. Our proposed technique consists of reverse-engineering the binary code running on bare-metal IoT devices and targeting the functions that control the PMIC configurations. We demonstrate our attack by modifying the firmware binary to alter the PMIC’s output voltage and evaluate it by measuring the changes in the output of the targeted sensors. We demonstrate that supplying a sensor with an incorrect voltage or current configuration can cause data corruption, which can go unnoticed and might have direct repercussions on real-world systems. Moreover, we discuss the stealthy nature of our attack and the fact that it can evade detection during functional testing as it does not change the overall functionality of IoT devices. Finally, we provide potential mitigation suggestions to address this vulnerability.
Sutej Kulkarni, Ryan Tsang, Asmita 0001, Houman Homayoun, Soheil Salehi
ICCD5
2023 Quantized Transformer Language Model Implementations on Edge Devices
abstract
Large-scale transformer-based models like the Bidi-rectional Encoder Representations from Transformers (BERT) are widely used for Natural Language Processing (NLP) applications, wherein these models are initially pre-trained with a large corpus with millions of parameters and then fine-tuned for a downstream NLP task. One of the major limitations of these large-scale models is that they cannot be deployed on resource- constrained devices due to their large model size and increased inference latency. In order to overcome these limitations, such large-scale models can be converted to an optimized FlatBuffer format, tailored for deployment on resource-constrained edge devices. Herein, we evaluate the performance of such FlatBuffer transformed MobileBERT models on three different edge devices, fine-tuned for Reputation analysis of English language tweets in the Rep Lab 2013 dataset. In addition, this study encompassed an evaluation of the deployed models, wherein their latency, performance, and resource efficiency were meticulously assessed. Our experiment results show that, compared to the original BERT large model, the converted and quantized MobileBERT models have 160x smaller footprints for a 4.1 % drop in accuracy while analyzing at least one tweet per second on edge devices. Furthermore, our study highlights the privacy-preserving aspect of TinyML systems as all data is processed locally within a serverless environment.
Mohammad Wali Ur Rahman, Murad Mehrab Abrar, Hunter Gibbons Copening, Salim Hariri, Sicong Shao, Pratik Satam, Soheil Salehi
ICMLA7
2023 Energy-/Area-Efficient Spintronic ANN-based Digit Recognition via Progressive Modular Redundancy
abstract
Neural networks offer viable alternatives for energy versus accuracy tradeoffs, in particular with regards to the precision of the computational circuit. This paper explores use of progressive modular redundancy of intrinsically low energy, low precision circuits as an alternative to more complex networks yielding higher accuracy directly. Results indicate that a lower footprint temporal modular redundancy, which is applied progressively as needed, can have lower footprint and reduced energy consumption at comparable or slightly reduced accuracy as more complex neural networks. This provides an alternative to binarization and other model compression options for intelligence at the edge of the network. Our Progressive Modular Redundancy approach using varied activations implemented using a$\mathbf{784}\times \mathbf{100}\times \mathbf{10}$network shows a 3% improvement in accuracy compared to the baseline case of$\mathbf{784}\times \mathbf{500}\times \mathbf{500}\times \mathbf{10}$network with sigmoidal activation, at 86.1% and 87% reduction in power and weighted crossbar normalized area overhead, respectively, 87.5% reduction in power error product (PEP) at the cost of ~2.6x increased throughput latency.
Mousam Hossain, Adrian Tatulian, Harshavardhan Reddy Thummala, Ronald F. DeMara, Soheil Salehi
ISCAS5
2023 Hardware Trojan Detection Using Machine Learning: A Tutorial
abstract
With the growth and globalization of IC design and development, there is an increase in the number of Designers and Design houses. As setting up a fabrication facility may easily cost upwards of $20 billion, costs for advanced nodes may be even greater. IC design houses that cannot produce their chips in-house have no option but to use external foundries that are often in other countries. Establishing trust with these external foundries can be a challenge, and these foundries are assumed to be untrusted. The use of these untrusted foundries in the global semiconductor supply chain has raised concerns about the security of the fabricated ICs targeted for sensitive applications. One of these security threats is the adversarial infestation of fabricated ICs with a Hardware Trojan (HT) . An HT can be broadly described as a malicious modification to a circuit to control, modify, disable, or monitor its logic. Conventional VLSI manufacturing tests and verification methods fail to detect HT due to the different and un-modeled nature of these malicious modifications. Current state-of-the-art HT detection methods utilize statistical analysis of various side-channel information collected from ICs, such as power analysis, power supply transient analysis, regional supply current analysis, temperature analysis, wireless transmission power analysis, and delay analysis. To detect HTs, most methods require a Trojan-free reference golden IC. A signature from these golden ICs is extracted and used to detect ICs with HTs. However, access to a golden IC is not always feasible. Thus, a mechanism for HT detection is sought that does not require the golden IC. Machine Learning (ML) approaches have emerged to be extremely useful in helping eliminate the need for a golden IC. Recent works on utilizing ML for HT detection have been shown to be promising in achieving this goal. Thus, in this tutorial, we will explain utilizing ML as a solution to the challenge of HT detection. Additionally, we will describe the Electronic Design Automation (EDA) tool flow for automating ML-assisted HT detection. Moreover, to further discuss the benefits of ML-assisted HT detection solutions, we will demonstrate a Neural Network (NN) -assisted timing profiling method for HT detection. Finally, we will discuss the shortcomings and open challenges of ML-assisted HT detection methods.
Kevin Immanuel Gubbi, Banafsheh S. Latibari, Anirudh Srikanth, Tyler David Sheaves, Sayed Arash Beheshti, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Avesta Sasan, Houman Homayoun, Soheil Salehi
ACM Trans. Embed. Comput. Syst.10
2022 LOCK&ROLL: deep-learning power side-channel attack mitigation using emerging reconfigurable devices and logic locking
abstract
The security and trustworthiness of ICs are exacerbated by the modern globalized semiconductor business model. This model involves many steps performed at multiple locations by different providers and integrates various Intellectual Properties (IPs) from several vendors for faster time-to-market and cheaper fabrication costs. Many existing works have focused on mitigating the well-known SAT attack and its derivatives. Power Side-Channel Attacks (PSCAs) can retrieve the sensitive contents of the IP and can be leveraged to find the key to unlock the obfuscated circuit without simulating powerful SAT attacks. To mitigate P-SCA and SAT-attack together, we propose a multi-layer defense mechanism called LOCK&ROLL: Deep-Learning Power Side-Channel Attack Mitigation using Emerging Reconfigurable Devices and Logic Locking. LOCK&ROLL utilizes our proposed Magnetic Random-Access Memory (MRAM)-based Look Up Table called Symmetrical MRAM-LUT (SyM-LUT). Our simulation results using 45nm technology demonstrate that the SyM-LUT incurs a small overhead compared to traditional Static Random Access Memory LUT (SRAM-LUT). Additionally, SyM-LUT has a standby energy consumption of 20aJ while consuming 33fJ and 4.6fJ for write and read operations, respectively. LOCK&ROLL is resilient against various attacks such as SAT-attacks, removal attack, scan and shift attacks, and P-SCA.
Gaurav Kolhe, Tyler David Sheaves, Kevin Immanuel Gubbi, Soheil Salehi, Setareh Rafatirad, Sai Manoj Pudukotai Dinakarrao, Avesta Sasan, Houman Homayoun
DAC4
2022 Survey of Machine Learning for Electronic Design Automation
abstract
An increase in demand for semiconductor ICs, recent advancements in machine learning, and the slowing down of Moore's law have all contributed to the increased interest in using Machine Learning (ML) to enhance Electronic Design Automation (EDA) and Computer-Aided Design (CAD) tools and processes. This paper provides a comprehensive survey of available EDA and CAD tools, methods, processes, and techniques for Integrated Circuits (ICs) that use machine learning algorithms. The ML-based EDA/CAD tools are classified based on the IC design steps. They are utilized in Synthesis, Physical Design (Floorplanning, Placement, Clock Tree Synthesis, Routing), IR drop analysis, Static Timing Analysis (STA), Design for Test (DFT), Power Delivery Network analysis, and Sign-off. The current landscape of ML-based VLSI-CAD tools, current trends, and future perspectives of ML in VLSI-CAD are also discussed.
Kevin Immanuel Gubbi, Sayed Aresh Beheshti-Shirazi, Tyler David Sheaves, Soheil Salehi, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
ACM Great Lakes Symposium on VLSI4
2022 FANDEMIC: Firmware Attack Construction and Deployment on Power Management Integrated Circuit and Impacts on IoT Applications
Ryan Tsang, Doreen Joseph, Asmita 0001, Soheil Salehi, Nadir Carreon, Prasant Mohapatra, Houman Homayoun
NDSS4
2021 Securing Hardware via Dynamic Obfuscation Utilizing Reconfigurable Interconnect and Logic Blocks
abstract
Maximizing profits while minimizing risk in a technologically advanced silicon industry has motivated the globalization of the fabrication process and electronic hardware supply chain. However, with the increasing magnitude of successful hardware attacks, the security of many hardware IPs has been compromised. Many existing security works have focused on resolving a single vulnerability while neglecting other threats. This motivated to propose a novel approach for securing hardware IPs during the fabrication process and supply chain via logic obfuscation by utilizing emerging spin-based devices. Our proposed dynamic obfuscation approach uses reconfigurable logic and interconnects blocks (RIL-Blocks), consisting of Magnetic Random Access Memory (MRAM)-based Look Up Tables and switch boxes flexibility and resiliency against state-of-the-art SAT-based attacks and power side-channel attacks while incurring a small overhead. The proposed Scan Enabled Obfuscation circuitry obfuscates the oracle circuit’s responses and further fortifies the logic and routing obfuscation provided by the RIL-Blocks, resembling a defense-in-depth approach. The empirical evaluation of security provided by the proposed RIL-Blocks on the ISCAS benchmark and common evaluation platform (CEP) circuit shows that resiliency comes with reduced overhead while providing resiliency to various hardware security threats.
Gaurav Kolhe, Soheil Salehi, Tyler David Sheaves, Houman Homayoun, Setareh Rafatirad, Sai Manoj Pudukotai Dinakarrao, Avesta Sasan
DAC2
2019 Workshop on Virtualized Active Learning in STEM
abstract
Summary form only given. Virtualized Active Learning (VAL) engages synchronous group-based problem solving within the last 30minutes of fully-live class offerings or the Face-to-Face component of mixed-mode delivery. As opposed to students solving problems together on-paper, which is not readily observable by the instructor nor scalable to larger class enrollments, VAL deploys laptops/tablets and Wi-Fi connectivity to create a virtualized active learning environment where students and instructors can interact.
Ronald F. DeMara, Soheil Salehi
FIE2
2019 Virtualized Active Learning for Undergraduate Engineering Disciplines (VALUED): A Pilot in a Large Enrollment STEM Classroom
abstract
This student poster paper presents an innovative practice in order to increase the scalability and efficacy of student problem-based team learning in large enrollment engineering classrooms. We have devised a novel Virtualized Active Learning (VAL) approach to facilitate instructional delivery, assessment, and review of teams. VAL introduces a new pathway by utilizing open-source digital environments for effective and scalable team-based learning in classroom settings while empowering equitable participation from diverse learners. The proposed method provides a unique opportunity for learners to acquire knowledge and skills that are considered vital in STEM fields such as working in multidisciplinary teams proficiently and communicating with other team members in an effective manner. Meanwhile, they are engaged in finding an optimal solution to a design problem that requires certain specific constraints to be adequately met. The results of our pilot study indicate excellent potential for VAL in large enrollment STEM courses while facilitating the instructors to provide assistance and feedback to students in real-time.
Soheil Salehi, Ronald F. DeMara
FIE1
2019 Clockless Spin-based Look-Up Tables with Wide Read Margin
abstract
In this paper, we develop a 6-input fracturable non-volatile Clockless LUT (C-LUT) using spin Hall effect (SHE)-based Magnetic Tunnel Junctions (MTJs) and provide a detailed comparison between the SHE-MTJ-based C-LUT and Spin Transfer Torque (STT)-MTJ-based C-LUT. The proposed C-LUT offers an attractive alternative for implementing combinational logic as well as sequential logic versus previous spin-based LUT designs in the literature. Foremost, C-LUT eliminates the sense amplifier typically employed by using a differential polarity dual MTJ design, as opposed to a static reference resistance MTJ. This realizes a much wider read margin and the Monte Carlo simulation of the proposed fracturable C-LUT indicates no read and write errors in the presence of a variety of process variations scenarios involving MOS transistors as well as MTJs. Additionally, simulation results indicate that the proposed C-LUT reduces the standby power dissipation by 5.4-fold compared to the SRAM-based LUT. Furthermore, the proposed SHE-MTJ-based C-LUT reduces the area by 1.3-fold and 2-fold compared to the SRAM-based LUT and the STT-MTJ-based C-LUT, respectively.
Soheil Salehi, Ramtin Zand, Ronald F. DeMara
ACM Great Lakes Symposium on VLSI1
2019 AQuRate: MRAM-based Stochastic Oscillator for Adaptive Quantization Rate Sampling of Sparse Signals
abstract
Recently, the promising aspects of compressive sensing have inspired new circuit-level approaches for their efficient realization within the literature. However, most of these recent advances involving novel sampling techniques have been proposed without considering hardware and signal constraints. Additionally, traditional hardware designs for generating non-uniform sampling clock incur large area overhead and power dissipation. Herein, we propose a novel non-uniform clock generator called Adaptive Quantization Rate (AQR) generator using Magnetic Random Access Memory (MRAM)-based stochastic oscillator devices. Our proposed AQR generator provides ~25-fold reduction in area, on average, while offering ~6-fold reduced power dissipation, on average, compared to the state-of-the-art non-uniform clock generators.
Soheil Salehi, Ramtin Zand, Alireza Zaeemzadeh, Nazanin Rahnavard, Ronald F. DeMara
ACM Great Lakes Symposium on VLSI1
2019 Self-Organized Sub-bank SHE-MRAM-based LLC: An energy-efficient and variation-immune read and write architecture
Soheil Salehi, Navid Khoshavi, Ramtin Zand, Ronald F. DeMara
Integr.1
2018 BGIM: Bit-Grained Instant-on Memory Cell for Sleep Power Critical Mobile Applications
abstract
This paper devises a novel energy-aware Non-Volatile Static Random Access Memory (NV-SRAM) framework for sleep power critical mobile applications. The beyond-Complementary Metal Oxide Semiconductor (CMOS) hardware architecture has been designed to minimize the overall static and leakage energy consumption while providing fast back-up and restore operations. Differential Spin-Hall Effect Magnetic Random Access Memory devices are utilized to realize the proposed framework called Bit-Grained Instant-on Memory Cell (BGIM). Our results indicate that the proposed BGIM consumes 121.51fJ on average for each back-up operation and 1.56fJ on average for each restore operation. Furthermore, the proposed BGIM can perform rapid back-up operations in 1ns and fast restore operations in 13.2ps. Moreover, the proposed BGIM cell incurs 0.4μm^2 area overhead compared to the traditional 6T SRAM cell, however it eliminates the need for data transmission and a separate non-volatile memory macro.
Soheil Salehi, Ronald F. DeMara
ICCD1
2017 Process variation immune and energy aware sense amplifiers for resistive non-volatile memories
abstract
Spin-Transfer Torque Magnetic Random Access Memory (STT-MRAM) has been explored as a post-CMOS technology for embedded and data storage applications seeking non-volatility, near-zero standby energy, and high density. Towards attaining these objectives for practical implementations, various techniques to mitigate the specific reliability challenges associated with STT-MRAM elements are surveyed, classified, and assessed herein. Some solutions to the reliability issues identified are addressed to realize reliable STT-MRAM designs. In an attempt to further improve the process variation immunity of the Sense Amplifiers (SAs), two new SAs are introduced: Energy Aware Sense Amplifier (EASA) and Variation Immune Sense Amplifier (VISA). Results have shown that EASA and VISA achieve superior performance in most cases compared to two of the most common SAs, namely PCSA and SPCSA respectively, while reducing Bit Error Rate (BER) and increasing reliability.
Soheil Salehi, Ronald F. DeMara
ISCAS1
2017 Survey of STT-MRAM Cell Design Strategies: Taxonomy and Sense Amplifier Tradeoffs for Resiliency
abstract
Spin-Transfer Torque Random Access Memory (STT-MRAM) has been explored as a post-CMOS technology for embedded and data storage applications seeking non-volatility, near-zero standby energy, and high density. Towards attaining these objectives for practical implementations, various techniques to mitigate the specific reliability challenges associated with STT-MRAM elements are surveyed, classified, and assessed in this article. Cost and suitability metrics assessed include the area of nanomagmetic and CMOS components per bit, access time and complexity, sense margin, and energy or power consumption costs versus resiliency benefits. Solutions to the reliability issues identified are addressed within a taxonomy created to categorize the current and future approaches to reliable STT-MRAM designs. A variety of destructive and non-destructive sensing schemes are assessed for process variation tolerance, read disturbance reduction, sense margin, and write polarization asymmetry compensation. The highest resiliency strategies deliver a sensing margin above 300mV while incurring low power and energy consumption on the order of picojoules and microwatts, respectively, and attaining read sense latency of a few nanoseconds down to hundreds of picoseconds for non-destructive and destructive sensing schemes, respectively.
Soheil Salehi, Deliang Fan, Ronald F. DeMara
ACM J. Emerg. Technol. Comput. Syst.1
2015 Reactive rejuvenation of CMOS logic paths using self-activating voltage domains
abstract
Although the trend of technology scaling is sought to realize higher performance computer systems, it also results in Integrated Circuits (ICs) suffering from increasing Process, Voltage, and Temperature (PVT) variations and adverse aging effects. In most cases, these reliability threats manifest themselves as timing errors on critical speed-paths of the circuit, if a large design guardband is not reserved. In this work, we propose the Reactive Rejuvenation (RR) architectural approach consisting of detection and recovery phases to mitigate circuit from BTI-induced aging. The BTI impact on the critical and near critical paths performance is continuously examined through a lightweight logic circuit which asserts an error signal in the case of any timing violation in those paths. By utilizing timing violation occurrence in the system, the timing-sensitive portion of the circuit is recovered from BTI through switching computations to redundant aging-critical voltage domain. The proposed technique achieves aging mitigation and reduced energy consumption as compared to a baseline circuit. Thus, significant voltage guardbands to meet the desired timing specification are avoided.
Rizwan A. Ashraf, Ahmad Alzahrani 0001, Navid Khoshavi, Ramtin Zand, Soheil Salehi, Arman Roohi, Mingjie Lin, Ronald F. DeMara
ISCAS5