Setareh Rafatirad

dblp:99/7115 · also Setareh Rafatirah · DBLP profile ↗
← Back
89ranked-venue papers
0as first author
50since 2021 · last 2026
0000-0003-2035-8512ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 58 · 28 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 12 since 2021Software engineering, systems software and programming languages · 11 · 5 since 2021Artificial intelligence and machine learning · 5Security and privacy · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Kumo: A Security-Focused Serverless Cloud Simulator
Khaled N. Khasawneh, Setareh Rafatirad, Houman Homayoun, Chongzhou Fang
CCGrid3
2026 HALO: A Typed Multi-View Graph Abstraction for RTL and Netlist Learning
abstract
Graph learning is now a dominant tool for analyzing and optimizing hardware designs, yet its effectiveness is repeatedly bottlenecked by circuit representation instead of neural network. We present Halo, a typed multi-view heterogeneous-graph abstraction framework that unifies register-transfer level (RTL) and gate-level netlists. Halo explicitly represents syntax, dataflow, control, sequential/time, hierarchy, hyperedge-correct nets, and cross-level alignment between RTL objects and their synthesized counterparts. We developed deterministic conversion pipelines that maps elaborated SystemVerilog designs using Surelog/UHDM and synthesized netlists, extracted through Yosys JSON into PyTorch-Geometric and DGL heterographs under an append-only schema. HALO is evaluated on a compact representative set of downstream tasks, including cross-level alignment prediction, vulnerability detection, localization, and robustness analysis across synthesis flows.
Kevin Immanuel Gubbi, Mohammadnavid Tarighat, Brinda Puri, Mahdi Pirayesh Shirazi Nejad, Setareh Rafatirad, Houman Homayoun
ACM Great Lakes Symposium on VLSI5
2026 Lightweight Cross-Device Sleep Tracking on the WeBe Wearable Platform
abstract
Wearable devices are widely used for continuous health monitoring, yet reliable sleep tracking on emerging platforms remains underexplored due to reliance on proprietary algorithms and device-specific activity representations. We present a lightweight and reproducible sleep tracking pipeline that operates directly on raw accelerometer signals. The method converts data into epoch-level activity features, applies temporal smoothing and normalized scoring, and performs sleep/wake classification using a globally calibrated threshold. We calibrate the model on the Multilevel Monitoring of Activity and Sleep in Healthy People (MMASH) dataset and evaluate it in a cross-device study using the WeBe wearable platform and a commercial ActiGraph device. On MMASH, the method achieves a mean absolute error of 41.6 minutes in Total Sleep Time (TST), with onset and offset errors of 6.3 and 7.4 minutes. On real-world WeBe data from three participants across five sessions, it achieves a mean TST error of 27.4 minutes and onset and offset errors of 13.9 and 8.0 minutes. In contrast, a commercial ActiGraph pipeline shows larger discrepancies relative to ground truth. These results demonstrate accurate and generalizable sleep tracking using a simple and reproducible pipeline.
Ehsan Kourkchi, Krishi Prashant Shah, Zequan Liang, Setareh Rafatirad, Houman Homayoun
ACM Great Lakes Symposium on VLSI5
2026 Know Me by My Pulse: Toward Practical Continuous Authentication on Wearable Devices via Wrist-Worn PPG
Zequan Liang, Ruoyu Zhang 0002, Ruijie Fang, Ning Miao, Ehsan Kourkchi, Setareh Rafatirad, Houman Homayoun, Chongzhou Fang
NDSS7
2026 Bit of a Close Talker: A Practical Guide to Serverless Cloud Co-Location Attacks
Najmeh Nazari, Behnam Omidi, Setareh Rafatirad, Khaled N. Khasawneh, Houman Homayoun, Chongzhou Fang
NDSS4
2025 Rapid Adaptation of $\text{SpO}_{2}$ Estimation to Wearable Devices via Transfer Learning on Low-Sampling-Rate PPG
abstract
Blood oxygen saturation$(\text{SpO}_{2})$is a vital marker for healthcare monitoring. Traditional$\text{SpO}_{2}$estimation methods often rely on complex clinical calibration, making them unsuitable for low-power, wearable applications. In this paper, we propose a transfer learning-based framework for the rapid adaptation of$\mathrm{SpO}_{2}$estimation to energy-efficient wearable devices using low-sampling-rate (25Hz) dual-channel photoplethysmography (PPG). We first pretrain a bidirectional Long ShortTerm Memory (BiLSTM) model with self-attention on a public clinical dataset, then fine-tune it using data collected from our wearable We-Be band and an FDA-approved reference pulse oximeter. Experimental results show that our approach achieves a mean absolute error (MAE) of 2.967% on the public dataset and 2.624% on the private dataset, significantly outperforming traditional calibration and non-transferred machine learning baselines. Moreover, using 25Hz PPG reduces power consumption by 40% compared to 100 Hz, excluding baseline draw. Our method also attains an MAE of$3.284\%$in instantaneous$\mathrm{SpO}_{2}$prediction, effectively capturing rapid fluctuations. These results demonstrate the rapid adaptation of accurate, low-power$\mathrm{S p O}_{2}$monitoring on wearable devices without the need for clinical calibration.
Zequan Liang, Ruoyu Zhang 0002, Krishna Karthik, Ehsan Kourkchi, Setareh Rafatirad, Houman Homayoun
BSN6
2025 Generalizable Blood Pressure Estimation from Multi-Wavelength PPG Using Curriculum-Adversarial Learning
abstract
Accurate and generalizable blood pressure (BP) estimation is vital for the early detection and management of cardiovascular diseases. In this study, we enforce subjectlevel data splitting on a public multi-wavelength photoplethysmography (PPG) dataset and propose a generalizable BP estimation framework based on curriculum-adversarial learning. Our approach combines curriculum learning, which transitions from hypertension classification to BP regression, with domainadversarial training that confuses subject identity to encourage the learning of subject-invariant features. Experiments show that multi-channel fusion consistently outperforms single-channel models. On the four-wavelength PPG dataset, our method achieves strong performance under strict subject-level splitting, with mean absolute errors (MAE) of 14.2mmHg for systolic blood pressure (SBP) and 6.4mmHg for diastolic blood pressure (DBP). Additionally, ablation studies validate the effectiveness of both the curriculum and adversarial components. These results highlight the potential of leveraging complementary information in multi-wavelength PPG and curriculum-adversarial strategies for accurate and robust BP estimation.
Zequan Liang, Ruoyu Zhang 0002, Mahdi Pirayesh Shirazi Nejad, Ehsan Kourkchi, Setareh Rafatirad, Houman Homayoun
BSN6
2025 Self-Supervised and Topological Signal-Quality Assessment for Any PPG Device
abstract
Wearable photoplethysmography (PPG) is embedded in billions of devices, yet its optical waveform is easily corrupted by motion, perfusion loss, and ambient light—jeopardizing downstream cardiometric analytics. Existing signal-quality assessment (SQA) methods rely either on brittle heuristics or on data-hungry supervised models. We introduce the first fully unsupervised SQA pipeline for wrist PPG. Stage 1 trains a contrastive 1-D ResNet-18 on 276 h of raw, unlabeled data from heterogeneous sources (varying in device and sampling frequency), yielding optical-emitter- and motioninvariant embeddings (i.e., the learned representation is stable across differences in LED wavelength, drive intensity, and device optics, as well as wrist motion). Stage 2 converts each 512-D encoder embedding into a 4-D topological signature via persistent homology (PH) and clusters these signatures with HDBSCAN. To produce a binary signal-quality index (SQI), the acceptable PPG signals are represented by the densest cluster while the remaining clusters are assumed to mainly contain poor-quality PPG signals. Without re-tuning, the SQI attains Silhouette, Davies-Bouldin, and Calinski-Harabasz scores of$0.72,0.34$, and 6,173, respectively, on a stratified sample of 10,000 windows. In this study, we propose a hybrid self-supervised-learning-topological-dataanalysis (SSL-TDA) framework that offers a drop-in, scalable, cross-device quality gate for PPG signals.
Ruoyu Zhang 0002, Zequan Liang, Ehsan Kourkchi, Setareh Rafatirad, Houman Homayoun
BSN5
2025 FaRAccel: FPGA-Accelerated Defense Architecture for Efficient Bit-Flip Attack Resilience in Transformer Models
abstract
Forget and Rewire (FaR) methodology has demonstrated strong resilience against Bit-Flip Attacks (BFAs) on Transformer-based models by obfuscating critical parameters through dynamic rewiring of linear layers. However, the application of FaR introduces non-negligible performance and memory overheads, primarily due to the runtime modification of activation pathways and the lack of hardware-level optimization. To overcome these limitations, we propose FaRAccel, a novel hardware accelerator architecture implemented on FPGA, specifically designed to offload and optimize FaR operations. FaRAccel integrates reconfigurable logic for dynamic activation rerouting, and lightweight storage of rewiring configurations, enabling low-latency inference with minimal energy overhead. We evaluate FaRAccel across a suite of Transformer models and demonstrate substantial reductions in FaR inference latency and improvement in energy efficiency, while maintaining the robustness gains of the original FaR methodology. To the best of our knowledge, this is the first hardware-accelerated defense against BFAs in Transformers, effectively bridging the gap between algorithmic resilience and efficient deployment on real-world AI platforms.
Najmeh Nazari, Banafsheh S. Latibari, Elahe Hosseini, Fatemeh Movafagh, Chongzhou Fang, Hosein Mohammadi Makrani, Kevin Immanuel Gubbi, Abhijit Mahalanobis, Setareh Rafatirad, Hossein Sayadi, Houman Homayoun
ICCD9
2025 Extended Operational Life for Wearable Health Devices: A Hybrid TinyML and Server-Side ML Approach
abstract
Wearable devices equipped with sensors like ECG, PPG, and heart rate monitors are pivotal in health monitoring, yet they face significant challenges due to limited battery life, particularly when using complex machine learning (ML) models for continuous monitoring. This paper explores an innovative approach to enhance energy efficiency in wearable health monitoring applications by integrating TinyML with advanced ML models. We propose using energy-efficient TinyML models for initial binary classification to detect abnormalities in ECG signals. Upon detecting an anomaly, the data is transmitted to a server where more complex ML models, perform detailed multi-label classification to identify specific conditions. Our methodology leverages event-driven software frameworks and adaptive data collection strategies to optimize power consumption, extending the operational life of wearable devices without sacrificing accuracy. The proposed system is evaluated on several metrics, including accuracy, F1-score, energy consumption, and model latency. Our findings demonstrate that the integration of TinyML can significantly extend battery life while maintaining high levels of accuracy and reliability in health monitoring, presenting a promising solution for long-term wearable device deployment.
Najmeh Nazari, Vedant Patel, Chongzhou Fang, Setareh Rafatirad, Houman Homayoun
ISCAS4
2025 Large Language Models for Opioid-Induced Respiratory Depression Prediction in Hospitalized Patients: A Retrospective Study
abstract
Opioid-induced adverse events pose significant risks to hospitalized patients. However, there is a limited understanding of which patients on general care floors are at risk for Opioid-Induced Respiratory Depression (OIRD). This study aims to bridge that knowledge gap by utilizing the advancements of AI to interpret Electronic Medical Records (EMRs). Recently, Large Language Models (LLMs) have gained attention for their exceptional capabilities in understanding human language, which makes them crucial for AI systems in healthcare that focus on clinical narratives. In this study, we extracted 2,663 hospitalized adult patient records from UC Davis Medical Center archives between January 2010 and April 2020 to identify patients at high risk of OIRD. For this purpose, we employed clinical language models (BioBERT, ClinicalBERT, and GatorTron) and fine-tuned them on the OIRD dataset. Additionally, we leveraged the capabilities of GPT-4, a state-of-the-art LLM, to select the most informative risk factors and enhance the accuracy of the predictive models.
Elahe Hosseini, Abhinav Srinivas, Najmeh Nazari, Charity Hale, Setareh Rafatirad, Houman Homayoun
ACM Trans. Comput. Heal.5
2025 HWREx: AI-enabled Hardware Weakness and Risk Exploration and Storytelling Framework with LLM-assisted Mitigation Suggestion
abstract
The growing complexity of modern computing frameworks has led to an increase in cybersecurity vulnerabilities reported to the National Vulnerability Database (NVD). Extracting meaningful trends from this vast amount of unstructured data is challenging without proper tools and methodologies. Existing approaches lack a holistic strategy for vulnerability mitigation and prediction and effective knowledge extraction from the Common Weakness Enumeration (CWE), Common Vulnerability Exposure (CVE), and Common Attack Pattern Enumeration and Classification (CAPEC) databases. We introduce the AI-enabled Hardware Weakness and Risk Exploration and Storytelling Framework with LLM-assisted Mitigation Suggestion (HWREx), designed to address hardware vulnerabilities and IoT security. Our architecture features an Ontology-driven Storytelling capability that automates ontology updates to track vulnerability patterns and evolution over time, while offering mitigation strategies. It also clarifies the complex interrelations among CVEs, CWEs, and CAPECs through interactive visual knowledge graphs. Our framework achieved accuracy rates of 62% for CWE-CWE, 83% for CWE-CVE, and 77% for CWE-CAPEC linkage predictions. These graphs are instrumental for in-depth hardware weakness analysis and enable HWREx to deliver comprehensive assessments and actionable mitigation strategies. Additionally, HWREx utilizes Generative Pre-trained Transformers (GPT) to offer tailored mitigation suggestions.
Sujan Ghimire, Yu-Zheng Lin, Muntasir Mamun, Md Muhtasim Alam Chowdhury, Farhad Alemi, Shuyu Cai, Jinduo Guo, Banafsheh S. Latibari, Setareh Rafatirad, Pratik Satam, Soheil Salehi
ACM Trans. Design Autom. Electr. Syst.11
2025 Assessing and Mitigating Heterogeneity-Driven Security Threats in the Cloud
abstract
Cloud computing has become crucial for the commercial world due to its computational capacity, storage capabilities, scalability, software integration, and billing convenience. Initially, clouds were relatively homogeneous, but now diverse machine configurations in heterogeneous clouds are recognized for their improved application performance and energy efficiency. This shift is driven by the integration of various hardware to accommodate diverse user applications. However, alongside these advancements, security threats like micro-architectural attacks are increasing concerns for cloud providers and users. Studies like Repttack and Cloak & Co-locate highlight the vulnerability of heterogeneous clouds to co-location attacks, where attacker and victim instances are placed together. The ease of these attacks isn’t solely linked to heterogeneity but also correlates with how heterogeneous the target systems are. Despite this, no numerical metrics exist to quantify cloud heterogeneity. This article introduces the Heterogeneity Score (HeteroScore) to evaluate server setups and instances. HeteroScore significantly correlates with co-location attack security. The article also proposes strategies to balance diversity and security. This study pioneers the quantitative analysis connecting cloud heterogeneity and infrastructure security.
Chongzhou Fang, Najmeh Nazari, Behnam Omidi, Han Wang 0020, Aditya Puri, Manish Arora, Setareh Rafatirad, Houman Homayoun, Khaled N. Khasawneh
ACM Trans. Internet Techn.7
2024 Validation of WeBe Band During Physical Activities
abstract
Data reliability and algorithm robustness are both important for wearable devices. To validate the accuracy of a recently published research vehicle, WeBe band, we conducted a concurrent heart rate (HR) and galvanic skin response (GSR) validity study. WeBe band, Empatica E4 and MindWare, which is currently considered the gold standard for collecting these measures, are compared concurrently. Fifty healthy adult partic-ipants volunteered (female n=29, 49 in 18–25 age range, 1 in 26–30 age range; [mean (SD)]: height = 167.6 (8.9) cm, mass = 150.1 (33.1) lbs). Participants wore the WeBe band and the Empatica band on opposite wrists (alternating device placement between participants) and the MindWare electrodes were placed on the on chest, back, and palms. Each participant completed a study session (a total 51 minutes) that included sitting, standing, normal paced walking and faster paced walking. Data was processed and validity was measured though: mean absolute percent error (MAPE), Bland-Altman limits of aggreement (LOA) and concordance coefficient (rc). Results showed that WeBe band is valid under all conditions.
Ruijie Fang, Sally Hang, Ruoyu Zhang 0002, Chongzhou Fang, Setareh Rafatirad, Camelia E. Hostinar, Houman Homayoun
BSN5
2024 Architectural Whispers: Robust Machine Learning Models Fingerprinting via Frequency Throttling Side-Channels
abstract
Machine Learning (ML) security practices include hiding ML model architectures to protect intellectual property and prevent attacks. We introduce a novel fingerprinting attack using frequency throttling-based Side-Channel Attack (SCA) to detect an ML model's architecture family by converting power side-channel data into timing variations. This method involves using adversary kernels and a time series ML classifier to discern the architecture from execution time patterns during model operation. We achieved up to 96% accuracy in identifying known ML models' architecture families under Ring 0 privileges and we demonstrated its effectiveness across different platforms. Moreover, our code is publicly available 1.
Najmeh Nazari, Chongzhou Fang, Hosein Mohammadi Makrani, Behnam Omidi, Mahdi Eslamimehr, Setareh Rafatirad, Avesta Sasan, Hossein Sayadi, Khaled N. Khasawneh, Houman Homayoun
DAC6
2024 SpecScope: Automating Discovery of Exploitable Spectre Gadgets on Black-Box Microarchitectures
abstract
Transient execution attacks pose information leakage risks in current systems. Disabling speculative execution, though mitigating the issue, results in significant performance loss. Accurate identification of vulnerable gadgets is essential for balancing security and performance. However, uncovering all covert channels is challenging due to complex microarchitectural analysis. This paper introduces SpecScope, a framework for automating the detection of Spectre gadgets in code using a black-box microarchitecture approach. SpecScope focuses on contention between transient and non-transient instructions to precisely identify and reduce false-positive Spectre gadgets, minimizing mitigation overhead. Tested on public libraries, SpecScope outperforms existing methods, reducing False-Positive rates by 8.9% and increasing True-Positive rates by 10.4%.
Najmeh Nazari, Behnam Omidi, Chongzhou Fang, Hosein Mohammadi Makrani, Setareh Rafatirad, Avesta Sasan, Houman Homayoun, Khaled N. Khasawneh
DATE5
2024 The AI Companion in Education: Analyzing the Pedagogical Potential of ChatGPT in Computer Science and Engineering
abstract
Artificial Intelligence (AI), with ChatGPT as a prominent example, has recently taken center stage in various domains including higher education, particularly in Computer Science and Engineering (CSE). The AI revolution brings both convenience and controversy, offering substantial benefits while lacking formal guidance on their application. The primary objective of this work is to comprehensively analyze the pedagogical potential of ChatGPT in CSE education, understanding its strengths and limitations from the perspectives of educators and learners. We employ a systematic approach, creating a diverse range of educational practice problems within CSE field, focusing on various subjects such as data science, programming, AI, machine learning, networks, and more. According to our examinations, certain question types, like conceptual knowledge queries, typically do not pose significant challenges to ChatGPT, and thus, are excluded from our analysis. Alternatively, we focus our efforts on developing more in-depth and personalized questions and project-based tasks. These questions are presented to ChatGPT, followed by interactions to assess its effectiveness in delivering complete and meaningful responses. To this end, we propose a comprehensive five-factor reliability analysis framework to evaluate the responses. This assessment aims to identify when ChatGPT excels and when it faces challenges. Our study concludes with a correlation analysis, delving into the relationships among subjects, task types, and limiting factors. This analysis offers valuable insights to enhance ChatGPT's utility in CSE education, providing guidance to educators and students regarding its reliability and efficacy.
Zhangying He, Thomas Nguyen, Tahereh Miari, Mehrdad Aliasgari, Setareh Rafatirad, Hossein Sayadi
EDUCON5
2024 Interactive Framework for Cybersecurity Education and Future Workforce Development
abstract
This research-to-practice paper presents a novel pedagogical tool for hardware cybersecurity education and workforce development. The growing importance of hardware security has made it essential for individuals and organizations to understand hardware security principles and best practices. However, the current educational curriculum falls short of fulfilling these emerging demands due to the rapidly changing hardware security landscape and limited opportunities for hands-on training. To address these challenges, we propose and have developed the Interactive Hardware and Cybersecurity (I-HaC) Educational Framework, a pedagogical educational framework that supplements existing courses by leveraging generative AI for individualized instruction related to hardware and cybersecurity, data mining, and applied Machine Learning (ML), as well as data visualization to enhance cybersecurity education and workforce development. The framework is designed to be utilized by graduate and undergraduate Electrical and Computer Engineering (ECE) and Computer Science (CS) students for a comprehensive introduction to cybersecurity exploits and countermeasures in an interactive manner with hands-on components. Using I-HaC, we have developed tailored lab components for a diverse range of students and intend to release I-HaC as open-source for the benefit of the ECE and CS education community.
Sujan Ghimire, Md Muhtasim Alam Chowdhury, Ryan Tsang, Richard C. Yarnell, Emma Heckert, Jaeden Wolf Carpenter, Yu-Zheng Lin, Muntasir Mamun, Ronald F. DeMara, Setareh Rafatirad, Pratik Satam, Soheil Salehi
FIE10
2024 Securing On-Chip Learning: Navigating Vulnerabilities and Potential Safeguards in Spiking Neural Network Architectures
abstract
On-chip learning is the process of training or updating machine learning models directly on specialized hardware. This approach differs from traditional machine learning, which typically conducts training on external computing resources like Central Processing Units (CPUs) or Graphics Processing Units (GPUs). On-chip learning offers several advantages, including reduced latency, improved energy efficiency, enhanced privacy, and adaptability. Consequently, it holds great promise for enabling intelligent decision-making and adaptability in resource-constrained edge and IoT devices while addressing privacy concerns. In Spiking Neural Network (SNN), on-chip learning is enabled by adjusting synaptic weights, allowing the network’s behavior to dynamically align with desired outcomes. However, this adaptability may introduce potential security vulnerabilities. Unmitigated security risks in on-chip learning can lead to various threats, including data leaks, unauthorized access, and even adversarial manipulation of the learning process. This manuscript aims to provide a comprehensive overview of the security risks associated with on-chip learning, with a focus on potential vulnerabilities within the SNN architecture. We will explore real-world scenarios where these vulnerabilities can be exploited and outline protective measures and mitigation strategies to address these security concerns.
Najmeh Nazari, Kevin Immanuel Gubbi, Banafsheh S. Latibari, Md Muhtasim Alam Chowdhury, Chongzhou Fang, Avesta Sasan, Setareh Rafatirad, Houman Homayoun, Soheil Salehi
ISCAS7
2024 Forget and Rewire: Enhancing the Resilience of Transformer-based Models against Bit-Flip Attacks
Najmeh Nazari, Hosein Mohammadi Makrani, Chongzhou Fang, Hossein Sayadi, Setareh Rafatirad, Khaled N. Khasawneh, Houman Homayoun
USENIX Security Symposium5
2024 Optimized and Automated Secure IC Design Flow: A Defense-in-Depth Approach
abstract
The globalization of the manufacturing process and the supply chain for electronic hardware has been driven by the need to maximize profitability while lowering risk in a technologically advanced silicon sector. However, many hardware IPs’ security features have been broken because of the rise in successful hardware attacks. Existing security efforts frequently ignore numerous dangers in favor of fixing a particular vulnerability. This inspired the development of a unique method that uses emerging spin-based devices to obfuscate circuitry to secure hardware intellectual property (IP) during fabrication and the supply chain. We propose an Optimized and Automated Secure IC (OASIC) Design Flow, a defense-in-depth approach that can minimize overhead while maximizing security. Our EDA tool flow uses a dynamic obfuscation method that employs dynamic lockboxes, which include switch boxes and magnetic random access memory (MRAM)-based look-up tables (LUT) while offering minimal overhead and being flexible and resilient against modern SAT-based attacks and power side-channel attacks. An EDA tool flow for optimized lockbox insertion is also developed to generate SAT-resilient design netlists with the least power and area overhead. PPA metrics and security (SAT attack time) are provided to the designer for each lockbox insertion run. A verification methodology is provided to verify locked and unlocked designs for functional correctness. Finally, we use ISCAS’85 benchmarks to show that the EDA tool flow provides a secure hardware netlist with maximum security while considering power and area constraints. Our results indicate that the proposed OASIC design flow can maximize security while incurring less than 15% area overhead and maintaining a similar power footprint compared to the original design. OASIC design flow demonstrates improved performance as design size increases, which demonstrates the scalability of the proposed approach.
Kevin Immanuel Gubbi, Banafsheh S. Latibari, Md Muhtasim Alam Chowdhury, Afrooz Jalilzadeh, Erfan Yazdandoost Hamedani, Setareh Rafatirad, Avesta Sasan, Houman Homayoun, Soheil Salehi
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 Introducing an Open-Source Python Toolkit for Machine Learning Research in Physiological Signal based Affective Computing
abstract
In the realm of physiological-based affective computing, significant progress has been witnessed in machine learning over the last two decades. Nevertheless, the lack of consistency in measurement tools and data organization across diverse datasets poses a challenge when integrating new datasets for algorithm testing, research, and result comparison across multiple datasets. Despite the expansion of artificial intelligence-driven affective computing, a notable gap remains in the form of a comprehensive toolkit tailored for both machine learning researchers and psychologists who are new to the field of machine learning. In response to these challenges, we introduce a Python toolkit designed to fulfill two key roles: establishing a standardized benchmark for affective computing datasets and offering an all-encompassing toolkit for machine learning in physiological signal based affective computing. This toolkit encompasses vital components essential to the machine learning process, encompassing tasks like dataset integration and interpretation, signal preprocessing, feature derivation, post-processing, classification models, and evaluation metrics. Our proposed toolkit is designed for working with seven publicly available datasets, embracing five different modalities and incorporating twenty diverse machine learning models spanning from conventional options like the support vector machine (SVM) to cutting-edge deep learning models. To the best of our knowledge, the proposed toolkit stands as the pioneering initiative for creating a standardized dataset benchmarking system and a comprehensive solution tailored for machine learning applications in affective computing. The open-source codebase for the proposed toolkit is accessible via https://github.com/rjfang/pyAffeCT.
Ruijie Fang, Ruoyu Zhang 0002, Elahe Hosseini, Chongzhou Fang, Setareh Rafatirad, Houman Homayoun
BIBM5
2023 Emotion and Stress Recognition Utilizing Galvanic Skin Response and Wearable Technology: A Real-time Approach for Mental Health Care
abstract
In modern society, people are exposed to various stressors and negative emotions daily and they may cause mental and physical diseases such as depression, anxiety, high blood pressure, heart attacks, and stroke. Therefore, this paper delves into the potential of modern wearable technologies as a tool for real-time health monitoring. The advent of ubiquitous sensing has ushered in an era where physiological and behavioral measurements can be continuously recorded in daily life. One significant physiological marker is the Galvanic Skin Response (GSR), which exhibits noteworthy changes under different emotional states. We propose a machine learning-based emotion recognition framework. It includes a preprocessing stage that eliminates noise and extracts 87 features from the GSR data. To account for individual differences in physiological responses, we also introduce a novel normalization procedure per subject. Finally, a subset of dominant and discriminative features enhances the proposed framework’s performance. We conducted experiments on two datasets, the wearable stress and affect detection dataset (WESAD) for stress detection, and the multimodal MAHNOB-HCI dataset for emotion recognition. The results show that the Leave-One-Out method is capable of detecting stress with 97.03% accuracy. Moreover, the proposed method classifies arousal and valence with an accuracy of 82.20% and 82.57%, respectively.
Elahe Hosseini, Ruijie Fang, Ruoyu Zhang 0002, Setareh Rafatirad, Houman Homayoun
BIBM4
2023 Federated Learning with Heterogeneous Models for On-device Malware Detection in IoT Networks
abstract
IoT devices have been widely deployed in many applications to facilitate smart technology, increased portability, and seamless connectivity. Despite being widely adopted, security in IoT devices is often considered an afterthought due to resource and cost constraints. Among multiple security threats, malware attacks are observed to be a pivotal threat to IoT devices. Considering the spread of IoT devices and the threats they experience over time, deploying a static malware detector trained offline seems ineffective. On the other hand, on-device learning is an expensive or infeasible option due to the limited available resources on IoT devices. To overcome these challenges, this work employs ‘Federated Learning’ (FL) which enables timely updates to the malware detection models for increased security while mitigating the high communication or data storage overhead of centralized cloud approaches. Federated learning allows training machine learning models with decentralized data while preserving its privacy by design. However, one of the challenges with the FL is that the on-device models are required to be homogeneous, which may not be true in the case of networked IoT systems. As a panacea, we introduce a methodology to unify the models in the cloud with minimal overheads and an impact on on-device malware detection. We evaluate the proposed technique against homogeneous models in networked IoT systems encompassing Raspberry Pi devices. The experimental results and system efficiency analysis indicate that end-to-end training time is just 1.12× higher than traditional FL, testing latency is 1.63× faster, and malware detection performance is improved by 7% to 13% for resource-constrained IoT devices.
Sanket Shukla, Setareh Rafatirad, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
DATE2
2023 Towards Race and Gender Equity in Data Science Education
abstract
Data Science education is experiencing annual enrollment growth, driven in part by government initiatives aimed at promoting racial equity in STEM fields. It is vital to ensure that all students have access to the necessary learning resources. While gender bias in STEM has received significant attention, research on racial bias is relatively limited, and the combined effects of gender and racial biases in Data Science remain largely unexplored. The objective of this study is to investigate these biases in Data Science education by examining the preferences of 300 diverse students enrolled in Data Science related classes, with a specific focus on their preferences regarding teaching practices and methods. The impact of race and gender in Data Science education is an important area of study that has gained increased attention in recent years. Race and gender play a critical role in representation within the Data Science community. Stereotypes and biases can influence the experiences of individuals from different racial and gender backgrounds within Data Science education. Race and gender can impact data science pedagogy in various ways, influencing instructional practices, content delivery, and student experiences. Using machine learning techniques to understand biased pedagogy in data science education can be a valuable approach to uncover patterns and insights. We propose an effective unsupervised learning approach to examine the correlations between preferred pedagogical methods and gender-racial background of the participating students. Our findings show that student's gender has a significant correlation with learning style; e.g., female students prefer more private methods of interaction with TAs and instructors, such as anonymous polls and Zoom meetings. Furthermore, our study uncovers significant associations between race, gender, and learning styles, showing that Asian students exhibit higher levels of Classroom Interaction and Group Assignment Preference than their White counterparts. These findings are crucial for Data Science educators seeking to identify and address racial and gender biases, promoting a more inclusive and equitable learning environment, especially for underrepresented minorities. Recognizing these influences enables educators to implement targeted strategies for a supportive and engaging educational experience for all students.
Brendan Baird, Namya Radesh, Setareh Rafatirad, Hossein Sayadi
FIE3
2023 Unleashing the Potential of Reinforcement Learning for Enhanced Personalized Education
abstract
Providing personalized content that meets the diverse needs of learners is a significant challenge for today's learning systems. Current educational platforms lack the capability to effectively support the needs of learners with varying intellectual abilities, learning pace, preferences, and academic backgrounds. To overcome this challenge, there is a pressing need for sustainable educational tools that can adapt to the individual needs of student learners. Our study proposes Reinforced-EDU, an intelligent and effective reinforcement learning-guided framework based on Q-Learning technique. The framework adaptively schedules course assignments and educational activities based on students' characteristics, preferences, skills, and academic backgrounds. It considers different student-centered academic factors to prescribe a suitable learning plan that maximizes the students' overall grade and satisfaction rate while reducing dropouts. We identify four distinct learning modes tailored towards effective logical reasoning that align best with the students' preferences and characteristics. Our proposed intelligent method, Reinforced-EDU, can dynamically assign an optimal learning path with 95% accuracy based on students' preferences and pre-knowledge.
Chelsea William Fernandes, Tahereh Miari, Setareh Rafatirad, Hossein Sayadi
FIE3
2023 HeteroScore: Evaluating and Mitigating Cloud Security Threats Brought by Heterogeneity
Chongzhou Fang, Najmeh Nazari, Behnam Omidi, Han Wang 0020, Aditya Puri, Manish Arora, Setareh Rafatirad, Houman Homayoun, Khaled N. Khasawneh
NDSS7
2023 Hardware Trojan Detection Using Machine Learning: A Tutorial
abstract
With the growth and globalization of IC design and development, there is an increase in the number of Designers and Design houses. As setting up a fabrication facility may easily cost upwards of $20 billion, costs for advanced nodes may be even greater. IC design houses that cannot produce their chips in-house have no option but to use external foundries that are often in other countries. Establishing trust with these external foundries can be a challenge, and these foundries are assumed to be untrusted. The use of these untrusted foundries in the global semiconductor supply chain has raised concerns about the security of the fabricated ICs targeted for sensitive applications. One of these security threats is the adversarial infestation of fabricated ICs with a Hardware Trojan (HT) . An HT can be broadly described as a malicious modification to a circuit to control, modify, disable, or monitor its logic. Conventional VLSI manufacturing tests and verification methods fail to detect HT due to the different and un-modeled nature of these malicious modifications. Current state-of-the-art HT detection methods utilize statistical analysis of various side-channel information collected from ICs, such as power analysis, power supply transient analysis, regional supply current analysis, temperature analysis, wireless transmission power analysis, and delay analysis. To detect HTs, most methods require a Trojan-free reference golden IC. A signature from these golden ICs is extracted and used to detect ICs with HTs. However, access to a golden IC is not always feasible. Thus, a mechanism for HT detection is sought that does not require the golden IC. Machine Learning (ML) approaches have emerged to be extremely useful in helping eliminate the need for a golden IC. Recent works on utilizing ML for HT detection have been shown to be promising in achieving this goal. Thus, in this tutorial, we will explain utilizing ML as a solution to the challenge of HT detection. Additionally, we will describe the Electronic Design Automation (EDA) tool flow for automating ML-assisted HT detection. Moreover, to further discuss the benefits of ML-assisted HT detection solutions, we will demonstrate a Neural Network (NN) -assisted timing profiling method for HT detection. Finally, we will discuss the shortcomings and open challenges of ML-assisted HT detection methods.
Kevin Immanuel Gubbi, Banafsheh S. Latibari, Anirudh Srikanth, Tyler David Sheaves, Sayed Arash Beheshti, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Avesta Sasan, Houman Homayoun, Soheil Salehi
ACM Trans. Embed. Comput. Syst.7
2022 Towards Generalized ML Model in Automated Physiological Arousal Computing: A Transfer Learning-Based Domain Generalization Approach
abstract
Physiological signal-based pattern recognition has progressed significantly, such as automated pain assessment and stress detection. Public datasets provide a research platform to conduct machine learning studies. However, models trained from public datasets easily overfit that specific dataset and do not apply to unseen data collected in real-life scenarios. This paper proposes to use the transfer learning-based domain generalization technique to generalize the models to solve this issue. Data from different training domains are generalized, i.e., the dissimilarity is minimized by the proposed approach such that the model trained is generalized. We proved that the generalized model is more adaptive to new unseen data. Experiments have been done on the BioVid heat pain dataset and WESAD stress dataset, and results showed that our proposed methods significantly improve the model performance on new unseen data.
Ruijie Fang, Ruoyu Zhang 0002, Elahe Hosseini, Anna M. Parenteau, Sally Hang, Setareh Rafatirad, Camelia E. Hostinar, Mahdi Orooji, Houman Homayoun
BIBM6
2022 Prevent Over-fitting and Redundancy in Physiological Signal Analyses for Stress Detection
abstract
Stress detection is an emerging field. WESAD is a commonly used public dataset for automated stress detection. It contains physiological signals including ECG, EDA, EMG, ACC, BVP, EDA, and skin temperature. The time window approach is used to extract features from time-series physiological signals. We find in previous studies that a 60-second time window with a 0.25-second window shift is widely used, but such window settings may cause redundancy and over-fitting. Thus, we propose to use (1) new window settings and (2) normalization per subject to tackle this problem. The experiment results show that our proposed methods significantly increase the classification performance.
Ruijie Fang, Ruoyu Zhang 0002, Elahe Hosseini, Anna M. Parenteau, Sally Hang, Setareh Rafatirad, Camelia E. Hostinar, Mahdi Orooji, Houman Homayoun
BIBM6
2022 A Low Cost EDA-based Stress Detection Using Machine Learning
abstract
Stress is an inevitable part of our lives in modern society since in many situations people are exposed to various stressors daily. According to studies, long-term stress can cause mental and physical diseases such as depression, anxiety, high blood pressure, heart attacks, and stroke. Therefore, stress detection is one of the crucial areas of study to maintain a healthy life. Recently, by developing commercial wearable technologies, real-time and continuous data collection for personal stress monitoring becomes more feasible. Under stress conditions, there are notable changes in physiological signals such as heart rate, respiration, perspiration, and eye pupil dilation. Previous studies have shown that Electrodermal Activity (EDA), also known as Galvanic Skin Response (GSR), can identify stress. EDA measures changes in perspiration by detecting the changes in the electrical conductivity of the skin. This paper focuses on stress detection using only EDA wearable sensors and applied machine learning techniques. First, 87 different features are extracted from EDA signals. Then, the data are normalized per subject because of differences in individuals’ physiological responses. Finally, five dominant features in stress detection are selected. We used a publicly available dataset, namely, the wearable stress and affect detection dataset (WESAD) in this study. The results show that the One-Leave-Out method is capable of detecting stress with 97.03% accuracy.
Elahe Hosseini, Ruijie Fang, Ruoyu Zhang 0002, Anna M. Parenteau, Sally Hang, Setareh Rafatirad, Camelia E. Hostinar, Mahdi Orooji, Houman Homayoun
BIBM6
2022 Silicon validation of LUT-based logic-locked IP cores
abstract
Modern semiconductor manufacturing often leverages a fabless model in which design and fabrication are partitioned. This has led to a large body of work attempting to secure designs sent to an untrusted third party through obfuscation methods. On the other hand, efficient de-obfuscation attacks have been proposed, such as Boolean Satisfiability attacks (SAT attacks). However, there is a lack of frameworks to validate the security and functionality of obfuscated designs. Additionally, unconventional obfuscated design flows, which vary from one obfuscation to another, have been key impending factors in realizing logic locking as a mainstream approach for securing designs. In this work, we address these two issues for Lookup Table-based obfuscation. We study both Volatile and Non-volatile versions of LUT-based obfuscation and develop a framework to validate SAT runtime using machine learning. We can achieve unparallel SAT-resiliency using LUT-based obfuscation while incurring 7% area and less than 1% power overheads. Following this, we discuss and implement a validation flow for obfuscated designs. We then fabricate a chip consisting of several benchmark designs and a RISC-V CPU in TSMC 65nm for post functionality validation. We show that the design flow and SAT-runtime validation can easily integrate LUT-based obfuscation into existing CAD tools while adding minimal verification overhead. Finally, we justify SAT-resilient LUT-based obfuscation as a promising candidate for securing designs.
Gaurav Kolhe, Tyler David Sheaves, Kevin Immanuel Gubbi, Tejas Kadale, Setareh Rafatirad, Sai Manoj Pudukotai Dinakarrao, Avesta Sasan, Hamid Mahmoodi, Houman Homayoun
DAC5
2022 LOCK&ROLL: deep-learning power side-channel attack mitigation using emerging reconfigurable devices and logic locking
abstract
The security and trustworthiness of ICs are exacerbated by the modern globalized semiconductor business model. This model involves many steps performed at multiple locations by different providers and integrates various Intellectual Properties (IPs) from several vendors for faster time-to-market and cheaper fabrication costs. Many existing works have focused on mitigating the well-known SAT attack and its derivatives. Power Side-Channel Attacks (PSCAs) can retrieve the sensitive contents of the IP and can be leveraged to find the key to unlock the obfuscated circuit without simulating powerful SAT attacks. To mitigate P-SCA and SAT-attack together, we propose a multi-layer defense mechanism called LOCK&ROLL: Deep-Learning Power Side-Channel Attack Mitigation using Emerging Reconfigurable Devices and Logic Locking. LOCK&ROLL utilizes our proposed Magnetic Random-Access Memory (MRAM)-based Look Up Table called Symmetrical MRAM-LUT (SyM-LUT). Our simulation results using 45nm technology demonstrate that the SyM-LUT incurs a small overhead compared to traditional Static Random Access Memory LUT (SRAM-LUT). Additionally, SyM-LUT has a standby energy consumption of 20aJ while consuming 33fJ and 4.6fJ for write and read operations, respectively. LOCK&ROLL is resilient against various attacks such as SAT-attacks, removal attack, scan and shift attacks, and P-SCA.
Gaurav Kolhe, Tyler David Sheaves, Kevin Immanuel Gubbi, Soheil Salehi, Setareh Rafatirad, Sai Manoj Pudukotai Dinakarrao, Avesta Sasan, Houman Homayoun
DAC5
2022 CR-Spectre: Defense-Aware ROP Injected Code-Reuse Based Dynamic Spectre
abstract
Side-channel attacks have been a constant threat to computing systems. In recent times, vulnerabilities in the architecture were discovered and exploited to mount and execute a state-of-the-art attack such as Spectre. The Spectre attack exploits a vulnerability in the Intel-based processors to leak confidential data through the covert channel. There exist some defenses to mitigate the Spectre attack. Among multiple defenses, hardware-assisted attack/intrusion detection (HID) systems have received overwhelming response due to its low overhead and efficient attack detection. The HID systems deploy machine learning (ML) classifiers to perform anomaly detection to determine whether the system is under attack. For this purpose, a performance monitoring tool profiles the applications to record hardware performance counters (HPC), utilized for anomaly detection. Previous HID systems assume that the Spectre is executed as a standalone application. In contrast, we propose an attack that dynamically generates variations in the injected code to evade detection. The attack is injected into a benign application. In this manner, the attack conceals itself as a benign application and gen-erates perturbations to avoid detection. For the attack injection, we exploit a return-oriented programming (ROP)-based code-injection technique that reuses the code, called gadgets, present in the exploited victim's (host) memory to execute the attack, which, in our case, is the CR-Spectre attack to steal sensitive data from a target victim (target) application. Our work focuses on proposing a dynamic attack that can evade HID detection by injecting perturbations, and its dynamically generated variations thereof, under the cloak of a benign application. We evaluate the proposed attack on the MiBench suite as the host. From our experiments, the HID performance degrades from 90% to 16%, indicating our Spectre-CR attack avoids detection successfully.
Abhijitt Dhavlle, Setareh Rafatirad, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
DATE2
2022 Machine Learning to the Rescue: ML-Assisted Framework for Equity-Driven Education
abstract
Data Science has recently experienced a significant surge in producing undergraduate degree/certificate and course enrollments. This growth has resulted in straining program resources at many institutions and causing concern about how to most effectively respond to the rapidly growing demand and provide equal opportunities for all students to pursue their degree. The primary goal of this work is to provide an intelligent solution to automatically identify biases related to race and gender in the content of Data Science related courses such as Machine Learning (ML), Natural Language Processing (NLP), Data Mining, Data Analytics, and Knowledge Discovery just to name a few. While machine learning allows to build powerful predicting tools, it hasn’t been used sufficiently to detect biases rooted in gender and race in educational material. Hence, using effective machine learning techniques we aim to enhance the success rate of undergraduate students in Data Science education and programs while promoting engagement through identifying and mitigating potential gender and racial biases. To this end, we propose a novel bias detection approach which uses a combination of gender identification and sentiment classification for characterization of gender and race in educational contents. Our proposed framework can help educators and course developers to mitigate race and gender biases in their course materials and create an equitable learning experience for students of minority and under-represented groups.
Wooyoung Chung, Xiyu Zhang 0002, Zunaira Ahmad, Hossein Sayadi, Setareh Rafatirad
EDUCON5
2022 Survey of Machine Learning for Electronic Design Automation
abstract
An increase in demand for semiconductor ICs, recent advancements in machine learning, and the slowing down of Moore's law have all contributed to the increased interest in using Machine Learning (ML) to enhance Electronic Design Automation (EDA) and Computer-Aided Design (CAD) tools and processes. This paper provides a comprehensive survey of available EDA and CAD tools, methods, processes, and techniques for Integrated Circuits (ICs) that use machine learning algorithms. The ML-based EDA/CAD tools are classified based on the IC design steps. They are utilized in Synthesis, Physical Design (Floorplanning, Placement, Clock Tree Synthesis, Routing), IR drop analysis, Static Timing Analysis (STA), Design for Test (DFT), Power Delivery Network analysis, and Sign-off. The current landscape of ML-based VLSI-CAD tools, current trends, and future perspectives of ML in VLSI-CAD are also discussed.
Kevin Immanuel Gubbi, Sayed Aresh Beheshti-Shirazi, Tyler David Sheaves, Soheil Salehi, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
ACM Great Lakes Symposium on VLSI6
2022 RAFeL - Robust and Data-Aware Federated Learning-inspired Malware Detection in Internet-of-Things (IoT) Networks
abstract
Federated Learning (FL) is a decentralized machine learning in which the training data is distributed on the Internet-of-Things (IoT) devices and learns a shared global model by aggregating local updates. However, the training data can be poisoned and manipulated by malicious adversaries, contaminating locally computed updates. To prevent this, detecting malicious IoT devices is very important. Since the local updates are large because of the high volume of data, minimizing the communication overhead is also necessary. This paper proposes a "RAFeL" framework, comprising of two techniques to tackle the above issues, (1) a robust defense technique and (2) a "Performance-aware bit-wise encoding" technique. "Robust and Active Protection with Intelligent Defense (RAPID)" is a defense system that detects malicious IoT devices and restricts the participation of the contaminated local updates computed by these malicious devices. To minimize communication cost, "Performance-aware bit-wise encoding" selects the appropriate encoding scheme for individual split bits based on their significance and effect on FL performance. The results illustrate that the proposed framework shows a 1.2-1.8x higher compression rate than lossy and lossless encoding techniques and has an average accuracy drop of 3% to 10% even with a fraction of malicious devices.
Sanket Shukla, Gaurav Kolhe, Houman Homayoun, Setareh Rafatirad, Sai Manoj Pudukotai Dinakarrao
ACM Great Lakes Symposium on VLSI4
2022 Breakthrough to Adaptive and Cost-Aware Hardware-Assisted Zero-Day Malware Detection: A Reinforcement Learning-Based Approach
abstract
In this paper, we have identified and addressed pressing challenges associated with online and cost-effective malware detection based on Hardware Performance Counters (HPCs) information. Existing Hardware-Assisted Malware Detection (HMD) methods guided by standard Machine Learning (ML) algorithms have limited their study on detecting known signatures of malicious patterns; thus, neglecting to address unknown (zero-day) malware detection at run-time which is a more challenging problem since the malware HPC data does not match any known attack applications’ signatures in the existing database. In addition, prior works have not presented a flexible and balanced solution that considers the trade-off between detection rate and implementation cost for adaptive selection of the best performing ML algorithms for online malware detection. In this paper, we first propose a unified feature selection method based on a heterogeneous feature fusion technique to effectively determine the most important HPC events for low-cost yet accurate malware detection. Next, we present Reinforced-HMD, a novel reinforcement learning-based framework for adaptive and cost-aware hardware-assisted zero-day malware detection based on desired performance metric and available hardware resources. To this aim, six classical and two reinforcement learning algorithms are implemented and their efficiency is thoroughly analyzed for detecting unknown malware using HPC events. Experimental results demonstrate that our Reinforced-HMD framework based on Upper Confidence Bound (UCB) learning approach achieves an accurate and robust detection rate with a 96% in both F1-score and AUC metrics for flexible and efficient zero-day malware detection while utilizing an optimal set of built-in HPC events.
Zhangying He, Hosein Mohammadi Makrani, Setareh Rafatirad, Houman Homayoun, Hossein Sayadi
ICCD3
2022 Iron-Dome: Securing IoT Networked Systems at Runtime by Network and Device Characteristics to Confine Malware Epidemics
abstract
The rapid growth of IoT networks presents an enlarged "attack space" for the adversary and poses significant security risks on a large scale. A single device in a network that is compromised under the influence of a malware attack, has the potential to spread malware across the network. This leads to a plethora of attacks, including DoS and ceasing the network functionality. Given the scale of IoT networks and the connectivity among the devices, mere detection and quarantining of malware in IoT networks does not limit the propagation of malware in IoT networks. This work proposes an integrated defense, termed as "IRON-DOME", comprising of (1) an on-device application analyzer: Image-based Malware detector that utilizes grayscale images of executables, (2) Device dynamic behavior analysis: Reliable extraction and dynamic analysis of malware Hardware Performance Counter (HPC) values; and (3) Device communication trait analyzer: Uses network packet data analysis to confine and propagate malware in the IoT network. The proposed solution yields: (1) a runtime malware detection accuracy of 93% within 19 ns, (2) is resource and power efficient; it consumes 30% fewer resources and 40% less power than state-of-the art defense techniques.
Sanket Shukla, Abhijitt Dhavlle, Sai Manoj Pudukotai Dinakarrao, Houman Homayoun, Setareh Rafatirad
ICCD5
2022 Repttack: Exploiting Cloud Schedulers to Guide Co-Location Attacks
Chongzhou Fang, Han Wang 0020, Najmeh Nazari, Behnam Omidi, Avesta Sasan, Khaled N. Khasawneh, Setareh Rafatirad, Houman Homayoun
NDSS7
2022 Imitating Functional Operations for Mitigating Side-Channel Leakage
abstract
Inspired by the idiom, “Mitigation (prevention) is better than cure!”, this work presents a random yet cognitive side-channel mitigation technique that is independent of underlying architecture and/or operating system. Unlike malware and other cyber-attacks, side-channel attacks (SCAs) exploit the architectural and design vulnerabilities and obtain sensitive information through the side channels. In contrast to the existing randomization-based side-channel defenses, we introduce a cognitive perturbation-based defense, Covert-Enigma, where the introduced perturbations look legit, but lead to an incorrect observation when interpreted by the attacker. To achieve this, the perturbations are injected at appropriate time instances to introduce additional operations, thereby misleading the attacker making the extracted data futile. To further make the attack more intricate for the attacker, proposed Covert-Enigma offers two modes of operation, chosen by the user, to determine the kind of induced cognitive perturbations—arbitraryandcyclicmodes. Arbitrary mode selects a group of key bits and flips them during every execution of the victim. Cyclic mode exhibits similar behavior, except it selects a new set of bits to flip after “$N$” cycles as chosen by the user. The cognitive perturbations are introduced in the form of a wrapper application to the victim, thus imposing no requirements on architectural level modifications nor soft updates/edits to the operating system. We report rigorous evaluation of the proposed Covert-Enigma protecting RSA cryptosystem attacked by Flush+Reload crypto SCA along with the bit(s) recovered after observing RSA under attack. Compared to traditional randomization-based defenses, proposed cognitive Covert-Enigma leads to 50% less overhead.
Abhijitt Dhavlle, Setareh Rafatirad, Khaled N. Khasawneh, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 A Neural Network-Based Cognitive Obfuscation Toward Enhanced Logic Locking
abstract
Logic obfuscation is introduced as a pivotal defense against multiple hardware threats on integrated circuits (ICs), including reverse engineering (RE) and intellectual property (IP) theft. The effectiveness of logic obfuscation is challenged by recently introduced Boolean satisfiability (SAT) attack and its variants. A plethora of counter measures have also been proposed to thwart the SAT attack. Irrespective of the implemented defense against SAT attacks, large power, performance, and area overheads are seen to be indispensable. In contrast, we propose a cognitive solution, which is a neural network (NN)-based SAT-hard clause translator, SATConda, that incurs a minimal area and power overhead while preserving the original functionality with enhanced security. SATConda is incubated with a SAT-hard clause generator that translates the existing conjunctive normal form (CNF) through minimal perturbations, such as the inclusion of pair of inverters or buffers or adding new lightweight SAT-hard block depending on the provided CNF. For efficient SAT-hard clause generation, SATConda is equipped with a multilayer NN that first learns the dependencies of features (literals and clauses), followed by a long short-term memory (LSTM) network to validate and backpropagate the SAT-hardness for better learning and translation. Our proposed SATConda is evaluated on ISCAS’85 and ISCAS’89 benchmarks and is seen to successfully defend against multiple state-of-the-art SAT attacks devised for hardware RE. In addition, we also evaluate our proposed SATConda’s empirical performance against MiniSAT, Lingeling, and Glucose SAT solvers that form the base for numerous existing deobfuscation SAT attacks.
Rakibul Hassan, Gaurav Kolhe, Setareh Rafatirad, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Breaking the Design and Security Trade-off of Look-up-table-based Obfuscation
abstract
Logic locking and Integrated Circuit (IC) camouflaging are the most prevalent protection schemes that can thwart most hardware security threats. However, the state-of-the-art attacks, including Boolean Satisfiability (SAT) and approximation-based attacks, question the efficacy of the existing defense schemes. Recent obfuscation schemes have employed reconfigurable logic to secure designs against various hardware security threats. However, they have focused on specific design elements such as SAT hardness. Despite meeting the focused criterion such as security, obfuscation incurs additional overheads, which are not evaluated in the present works. This work provides an extensive analysis of Look-up-table (LUT)–based obfuscation by exploring several factors such as LUT technology, size, number of LUTs, and replacement strategy as they have a substantial influence on Power-Performance-Area (PPA) and Security (PPA/S) of the design. We show that using large LUT makes LUT-based obfuscation resilient to hardware security threats. However, it also results in enormous design overheads beyond practical limits. To make the reconfigurable logic obfuscation efficient in terms of design overheads, this work proposes a novel LUT architecture where the security provided by the proposed primitive is superior to that of the traditional LUT-based obfuscation. Moreover, we leverage the security-driven design flow, which uses off-the-shelf industrial EDA tools to mitigate the design overheads further while being non-disruptive to the current industrial physical design flow. We empirically evaluate the security of the LUTs against state-of-the-art obfuscation techniques in terms of design overheads and SAT-attack resiliency. Our findings show that the proposed primitive significantly reduces both area and power by a factor of 8 \( \times \) and 2 \( \times \) , respectively, without compromising security.
Gaurav Kolhe, Tyler David Sheaves, Sai Manoj Pudukotai Dinakarrao, Hamid Mahmoodi, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
ACM Trans. Design Autom. Electr. Syst.5
2021 Securing Hardware via Dynamic Obfuscation Utilizing Reconfigurable Interconnect and Logic Blocks
abstract
Maximizing profits while minimizing risk in a technologically advanced silicon industry has motivated the globalization of the fabrication process and electronic hardware supply chain. However, with the increasing magnitude of successful hardware attacks, the security of many hardware IPs has been compromised. Many existing security works have focused on resolving a single vulnerability while neglecting other threats. This motivated to propose a novel approach for securing hardware IPs during the fabrication process and supply chain via logic obfuscation by utilizing emerging spin-based devices. Our proposed dynamic obfuscation approach uses reconfigurable logic and interconnects blocks (RIL-Blocks), consisting of Magnetic Random Access Memory (MRAM)-based Look Up Tables and switch boxes flexibility and resiliency against state-of-the-art SAT-based attacks and power side-channel attacks while incurring a small overhead. The proposed Scan Enabled Obfuscation circuitry obfuscates the oracle circuit’s responses and further fortifies the logic and routing obfuscation provided by the RIL-Blocks, resembling a defense-in-depth approach. The empirical evaluation of security provided by the proposed RIL-Blocks on the ISCAS benchmark and common evaluation platform (CEP) circuit shows that resiliency comes with reduced overhead while providing resiliency to various hardware security threats.
Gaurav Kolhe, Soheil Salehi, Tyler David Sheaves, Houman Homayoun, Setareh Rafatirad, Sai Manoj Pudukotai Dinakarrao, Avesta Sasan
DAC5
2021 On-device Malware Detection using Performance-Aware and Robust Collaborative Learning
abstract
The proliferation of the Internet-of-Things (IoT) devices has facilitated smart connectivity and enhanced computational capabilities. Lack of proper security protocols in such devices makes them vulnerable to cyber threats, especially malware attacks. Given the diversity and sophistication in malware samples, detecting them using traditional vendor database-based signature matching techniques is inefficient. This paper presents a collaborative machine learning (ML)-based malware detection framework. We introduce a) performance-aware precision-scaled federated learning (FL) to minimize the communication overheads with minimal device-level computations; and (2) a Robust and Active Protection with Intelligent Defense strategy against malicious activity (RAPID) at the device and network-level due to malware and other cyber-attacks. Deploying FL facilitates detecting malware attacks through collaborative learning and prevents data sharing, thus ensuring data security and privacy. RAPID denies the illegitimate user and aids in developing an effective collaborative malware detection model. A comprehensive analysis, results, and performance of the proposed technique are presented along with the communication overheads. An average accuracy of 94% is obtained with the proposed technique with 15% communication overhead, indicating 19% better performance than state-of-the-art techniques. Furthermore, the minimum accuracy drop of a model trained using RAPID is only 3% when 10% of devices are adversarial and 16% even when 40% of devices are adversarial.
Sanket Shukla, Sai Manoj Pudukotai Dinakarrao, Gaurav Kolhe, Setareh Rafatirad
DAC4
2021 HMD-Hardener: Adversarially Robust and Efficient Hardware-Assisted Runtime Malware Detection
abstract
To overcome the performance overheads incurred by the traditional software-based malware detection techniques, machine learning (ML) based Hardware-assisted Malware Detection (HMD) has emerged as a panacea to detect malicious applications and provide security. HMD primarily relies on the generated low-level microarchitectural events captured through Hardware Performance Counters (HPCs). This work proposes an adversarial attack on the HMD systems to tamper the security by introducing perturbations in performance counter traces with an adversarial sample generator application. To craft the attack, we first deploy an adversarial sample predictor to predict the adversarial HPC pattern for a given application to be misclassified by the deployed ML classifier in the HMD. Further, as the attacker has no direct access to manipulate the HPCs generated during runtime, based on the adversarial sample predictor's output, devise an adversarial sample generator wrapped around the victim application to produce HPC patterns similar to the adversarial predictor's estimated trace. With the proposed attack, malware detection accuracy is reduced to 18.1% from 82%. To render the HMD robust to such attacks, we further propose adversarially training the HMD to demonstrate that hardening can render HMD resilient against attacks; the detection accuracy post hardening raises to 81.2%.
Abhijitt Dhavlle, Sanket Shukla, Setareh Rafatirad, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
DATE3
2021 A Cognitive SAT to SAT-Hard Clause Translation-based Logic Obfuscation
abstract
Logic obfuscation is introduced as a pivotal defense mechanism against emerging hardware threats on Integrated Circuits (ICs) such as reverse engineering (RE) and intellectual property (IP) theft. The effectiveness of logic obfuscation is challenged by recently introduced Boolean satisfiability (SAT) attack and it's variants. A plethora of counter measures have been proposed to thwart the SAT attacks. Irrespective of the implemented defenses, large power, performance and area (PPA) overheads are seen to be indispensable. In contrast, we propose a neural network-based cognitive SAT to SAT-hard clause translator under the constraints of minimal PPA overheads while preserving the original functionality with impenetrable security. Our proposed method is incubated with a SAT-hard clause generator that translates the existing conjunctive normal form (CNF) through minimal perturbations such as inclusion of pair of inverters or buffers or adding new lightweight SAT-hard block depending on the provided CNF. For efficient SAT-hard clause generation, the proposed method is equipped with a multi-layer neural network that first learns the dependencies of features (literals and clauses), followed by a long-short-term-memory (LSTM) network to validate and backpropagate the SAT-hardness for better learning and translation. For a fair comparison with the state-of-the-art, we evaluate our proposed technique on ISCAS'85 benchmarks. It is seen to successfully defend against multiple state-of-the-art SAT attacks devised for hardware RE. In addition, we also evaluate our proposed technique's empirical performance against MiniSAT, Lingeling and Glucose SAT solvers that form the base for numerous existing deobfuscation SAT attacks.
Rakibul Hassan, Gaurav Kolhe, Setareh Rafatirad, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
DATE3
2021 Energy-Efficient and Adversarially Robust Machine Learning with Selective Dynamic Band Filtering
abstract
The popularity of neural networks is increasing day by day. Traditional machine learning solutions, such as image recognition, object detection, are being replaced by deep learning solutions because of their vigorous performance in computer vision. Despite their superior performance in these applications, neural networks are prone to adversarial attacks. The adversarial attack is the process of using adversarial samples as an input to the neural network which causes the network to misclassify, eventually degrading overall performance. Thus, it becomes very important to maintain their robustness by identifying, analyzing, and eliminating the cause of their vulnerability. In this paper, we introduce a technique to determine the most sensitive frequency band of input samples and filter the noise from this band to shield the network against adversarial attacks. First, we decompose the input sample into four different frequency components and then, identify the sensitive component by measuring the change in behavior of the pre-trained network on normal frequency band and that on frequency band with added noise (frequency band of an adversary). Next, we exploit this vulnerable component to assist the network in tackling the adversaries through noise filtering. Thereby, enhancing the neural networks? performance and defending against the adversarial attack. The low-frequency component was the most vulnerable and mitigating the noise from this band significantly improved the accuracy of Convolutional Neural Networks (CNN) along with that of state-of-art networks against adversarial attacks such as Fast Gradient Sign Method (FGSM), DeepFool (DF), and other techniques. The proposed technique showed performance enhancement from 85% to 95% classification accuracy for ResNet50.
Neha Nagarkar, Khaled N. Khasawneh, Setareh Rafatirad, Avesta Sasan, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
ACM Great Lakes Symposium on VLSI3
2021 Performance-aware Malware Epidemic Confinement in Large-Scale IoT Networks
abstract
As millions of IoT devices are interconnected together for better communication and computation, compromising even a single device opens a gateway for the adversary to access the network leading to an epidemic. It is pivotal to detect any malicious activity on a device and mitigate the threat. Among multiple feasible security threats, malware (malicious applications) poses a serious risk to modern IoT networks. A wide range of malware can replicate itself and propagate through the network via the underlying connectivity in the IoT networks making the malware epidemic inevitable. There exist several techniques ranging from heuristics to game-theory based technique to model the malware propagation and minimize the impact on the overall network. The state-of-the-art game-theory based approaches solely focus either on the network performance or the malware confinement but does not optimize both simultaneously. In this paper, we propose a throughput-aware game theory-based end-to-end IoT network security framework to confine the malware epidemic while preserving the overall network performance. We propose a two-player game with one player being the attacker and other being the defender. Each player has three different strategies and each strategy leads to a certain gain to that player with an associated cost. A tailored min-max algorithm was introduced to solve the game. We have evaluated our strategy on a 500 node network for different classes of malware and compare with existing state-of-the-art heuristic and game theory-based solutions.
Rakibul Hassan, Setareh Rafatirad, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
ICC2
2021 HosNa: A DPC++ Benchmark Suite for Heterogeneous Architectures
abstract
Most data centers equipped their general-purpose processors with hardware accelerators to reduce power consumption and improve utilization. Hardware accelerators offer highly energy-efficient computation for a wide range of applications; however, their programming is not as efficient as processors. To bridge the gap, Intel developed a cloud-based infrastructure called DevCloud that connects Intel® Xeon® Scalable Processors to GPUs and FPGAs to deliver high compute performance for emerging workloads. DevCloud assists developers with their compute-intensive tasks and provides access to precompiled software optimized for Intel® architecture. To reduce programming complexity and minimize the barriers to adopt new innovative hardware technology, Intel also provided a unified, cross-architecture programming model called oneAPI based on the Data-Parallel C++ (DPC++) language. In this paper, we introduce HosNa, the first DPC++ benchmark suite that can be used for the evaluation of the Intel FPGAs and DPC++ productivity. Moreover, we present the characterization of proposed benchmarks and the evaluation of implemented hardware accelerators in terms of speedup and latency.
Najmeh Nazari, Hosein Mohammadi Makrani, Hossein Sayadi, Lawrence Landis, Setareh Rafatirad, Houman Homayoun
ICCD5
2020 Estimating the Circuit De-obfuscation Runtime based on Graph Deep Learning
abstract
Circuit obfuscation has been proposed to protect digital integrated circuits (ICs) from different security threats such as reverse engineering by introducing ambiguity in the circuit, i.e., the addition of the logic gates whose functionality cannot be determined easily by the attacker. In order to conquer such defenses, techniques such as Boolean satisfiability-checking (SAT)-based attacks were introduced. SAT-attack can potentially decrypt the obfuscated circuits. However, the deobfuscation runtime could have a large span ranging from few milliseconds to a few years or more, depending on the number and location of obfuscated gates, the topology of the obfuscated circuit and obfuscation technique used. To ensure the security of the deployed obfuscation mechanism, it is essential to accurately pre-estimate the deobfuscation time. Thereby one can optimize the deployed defense in order to maximize the deobfuscation runtime. However, estimating the deobfuscation runtime is a challenging task due to 1) the complexity and heterogeneity of the graph-structured circuit, 2) the unknown and sophisticated mechanisms of the attackers for deobfuscation, 3) efficiency and scalability requirement in practice. To address the challenges mentioned above, this work proposes the first machine-learning framework that predicts the deobfuscation runtime based on graph deep learning. Specifically, we design a new model, ICNet with new input and convolution layers to characterize the circuit's topology, which is then integrated by composite deep fully-connected layers to obtain the deobfuscation runtime. The proposed ICNet is an end-to-end framework that can automatically extract the deter-minant features required for deobfuscation runtime prediction. Extensive experiments on standard benchmarks demonstrate its effectiveness and efficiency beyond many competitive baselines.
Zhiqian Chen, Gaurav Kolhe, Setareh Rafatirad, Chang-Tien Lu, Sai Manoj Pudukotai Dinakarrao, Houman Homayoun, Liang Zhao 0002
DATE3
2020 Mitigating Cache-Based Side-Channel Attacks through Randomization: A Comprehensive System and Architecture Level Analysis
abstract
Cache hierarchy was designed to allow CPU cores to process instructions faster by bridging the significant latency gap between the main memory and processor. In addition, various cache replacement algorithms are proposed to predict future data and instructions to boost the performance of the computer systems. However, recently proposed cache-based Side-Channel Attacks (SCAs) have shown to effectively exploiting such a hierarchical cache design. The cache-based SCAs are exploiting the hardware vulnerabilities to steal secret information from users by observing cache access patterns of cryptographic applications and thus are emerging as a serious threat to the security of the computer systems. Prior works on mitigating the cache-based SCAs have mainly focused on cache partitioning techniques and/or randomization of mapping between main memory. However, such solutions though effective, require modification in the processor hardware which increases the complexity of architecture design and are not applicable to current as well as legacy architectures. In response, this paper proposes a lightweight system and architecture level randomization technique to effectively mitigate the impact of side-channel attacks on last-level caches with no hardware redesign overhead for current as well as legacy architectures. To this aim, by carefully adapting the processor frequency and prefetchers operation and adding proper level of noise to the attackers' cache observations we attempt to protect the critical information from being leaked. The experimental results indicate that the concurrent randomization of frequency and prefetchers can significantly prevent cache-based side-channel attacks with no need for a new cache design. In addition, the proposed randomization and adaptation methodology outperforms the stat-of-the-art solutions in terms of the performance and execution time by reducing the performance overhead from 32.66% to nearly 20%.
Han Wang 0020, Hossein Sayadi, Tinoosh Mohsenin, Liang Zhao 0002, Avesta Sasan, Setareh Rafatirad, Houman Homayoun
DATE6
2020 StealthMiner: Specialized Time Series Machine Learning for Run-Time Stealthy Malware Detection based on Microarchitectural Features
abstract
Hardware-Assisted Malware Detection (HMD) techniques deploy Machine Learning (ML) classifiers to detect patterns of malicious applications based on microarchitectural features captured by modern microprocessors' Hardware Performance Counters (HPCs). Existing HMD methods have limited their analysis on detecting malicious applications that are spawned as a separate thread during application execution, hence detecting embedded malware patterns at run-time still remains an important challenge. Embedded malware refers to harmful stealthy cyber attacks in which the malicious code is hidden within benign applications and remains undetected by traditional malware detection approaches. In HMD methods, when the HPC data is directly fed into a machine learning classifier, embedding malicious code inside the benign applications leads to contamination of HPC information, as the collected HPC features combine benign and malware microarchitectural events together. To address this challenge, in this paper we propose StealthMiner, a specialized time series machine learning approach to accurately detect embedded malware at run-time using branch instructions feature, the most prominent microarchitectural feature. The results indicate that StealthMiner can detect embedded malware at run-time with 94% detection performance on average with only one HPC feature, outperforming the detection performance of state-of-the-art HMD methods by 42%.
Hossein Sayadi, Yifeng Gao 0001, Hosein Mohammadi Makrani, Tinoosh Mohsenin, Avesta Sasan, Setareh Rafatirad, Jessica Lin 0001, Houman Homayoun
ACM Great Lakes Symposium on VLSI6
2020 Comprehensive Evaluation of Machine Learning Countermeasures for Detecting Microarchitectural Side-Channel Attacks
abstract
Microarchitectural Side-Channel Attacks (SCAs) have posed serious threats to the security of modern computing systems. Such attacks exploit side-channel vulnerabilities stemming from fundamental performance-enhancing components such as cache memories. The existing works on detection of SCAs based on low-level microarchitectural features have considered collecting both victim and attack applications' hardware events that are captured from processors' hardware performance counter (HPC) registers. However, in such techniques the attack HPCs data can be easily manipulated and/or corrupted resulting in misleading the SCAs detection mechanism. In addition, the prior studies have explored the suitability of a limited number of Machine Learning (ML) algorithms in detecting microarchitectural SCAs. In response, in this paper, we conduct a comprehensive evaluation of various machine learning-based countermeasures for real-time side-channel attack detection based on low-level microarchitectural features. For this purpose, the victim applications' behavior is collected using the HPC features and analyzed under no attack and attack conditions to avoid potential manipulation of attackers' HPCs. We further explore the HPCs monitoring overhead when microarchitectural features are sampled at different intervals to find out the appropriate sampling interval for SCAs detection. For the purpose of thorough analysis, various types of ML classifiers are implemented and precisely compared across different evaluation metrics including detection accuracy, F-measure, robustness (Area Under the ROC Curve), and computational latency to identify the most efficient ML classifiers for real-time microarchitectural SCAs detection
Han Wang 0020, Hossein Sayadi, Avesta Sasan, Setareh Rafatirad, Tinoosh Mohsenin, Houman Homayoun
ACM Great Lakes Symposium on VLSI4
2020 Hybrid-Shield: Accurate and Efficient Cross-Layer Countermeasure for Run-Time Detection and Mitigation of Cache-Based Side-Channel Attacks
abstract
Cache-based Side-Channel Attacks (SCAs) exploit the emerging hardware vulnerabilities to steal secret information by observing cache access patterns of cryptographic applications. To address the challenges introduced by SCAs, existing solutions either rely on detecting them using the profiled hardware-related information of victim and attack programs, or mitigating the cache-based SCAs by focusing on cache partitioning techniques and/or randomization of cache mappings by modifying the underlying cache architecture of the processor. However, the former's detectors highly rely on the knowledge of attack programs that are not always available or could be obfuscated in real-world benign applications resulting in misidentifying attacks. On the other hand, the latter approach though effective, requires modification in the processor hardware which increases the complexity of architecture design and are not applicable to current as well as legacy architectures. To address the drawbacks, we propose Hybrid-Shield, an accurate and efficient cross-layer countermeasure for run-time detection and mitigation of cache-based side-channel attacks. For the detection stage, microarchitectural information of victim under attack and under no attack conditions are collected for training machine learning classifiers. For the mitigation stage, Hybrid-Shield adapts hardware prefetchers and scales processor frequency to increase the noise level in observed cache access pattern attacks to induce secret information. The experimental results indicate that Hybrid-Shield can achieve 100% detection rate with 0% false alarm rate and detected attacks' error rate increases from less than 5% to above 35% with only 15% performance overhead.
Han Wang 0020, Hossein Sayadi, Avesta Sasan, Setareh Rafatirad, Houman Homayoun
ICCAD4
2020 Phased-Guard: Multi-Phase Machine Learning Framework for Detection and Identification of Zero-Day Microarchitectural Side-Channel Attacks
abstract
Microarchitectural Side-Channel Attacks (SCAs) have emerged recently to compromise the security of computer systems by exploiting the existing processors' hardware vulnerabilities. In order to detect such attacks, prior studies have proposed the deployment of low-level features captured from built-in Hardware Performance Counter (HPC) registers in modern microprocessors to implement accurate Machine Learning (ML)-based SCAs detectors. Though effective, such attack detection techniques have mainly focused on binary classification models offering limited insights on identifying the type of attacks. In addition, while existing SCAs detectors required prior knowledge of attacks applications to detect the pattern of side-channel attacks using a variety of microarchitectural features, detecting unknown (zero-day) SCAs at run-time using the available HPCs remains a major challenge. In response, in this work we first identify the most important HPC features for SCA detection using an effective feature reduction method. Next, we propose Phased-Guard, a two-level machine learning-based framework to accurately detect and classify both known and unknown attacks at run-time using the most prominent low-level features. In the first level (SCA Detection), Phased-Guard using a binary classification model detects the existence of SCAs on the target system by determining the critical scenarios including system under attack and system under no attack. In the second level (SCA Identification) to further enhance the security against side-channel attacks, Phased-Guard deploys a multiclass classification model to identify the type of SCA applications. The experimental results indicate that Phased-Guard by monitoring only the victim applications' microarchitectural HPCs data, achieves up to 98 % attack detection accuracy and 99.5% SCA identification accuracy significantly outperforming the state-of-the-art solutions by up to 82 % in zero-day attack detection at the cost of only 4% performance overhead for monitoring.
Han Wang 0020, Hossein Sayadi, Gaurav Kolhe, Avesta Sasan, Setareh Rafatirad, Houman Homayoun
ICCD5
2020 HybriDG: Hybrid Dynamic Time Warping and Gaussian Distribution Model for Detecting Emerging Zero-Day Microarchitectural Side-Channel Attacks
abstract
Microarchitectural Side-channel Attacks (SCAs) benefit from emerging hardware vulnerabilities in modern microprocessors to steal critical information from users, posing great security threats to computer systems. Several recent studies have focused on using low-level features captured from built-in Hardware Performance Counter (HPC) registers to implement accurate Machine Learning (ML)-based SCAs detectors. Nonetheless, existing ML-based SCAs detectors required prior knowledge of attacks applications to detect the pattern of side-channel attacks using a variety of microarchitectural features. In particular, the existing solutions have ignored to address the challenge of detecting sophisticated unknown (zero-day) SCAs at run-time which is a more challenging issue in today's computer systems. In addition, prior works analyzed a limited number of ML classifiers without thoroughly evaluating the detection effectiveness and computational complexity of the detectors. In response, we propose HybriDG, a hybrid lightweight model consisting of Dynamic Time Warping (DTW) followed by a Gaussian distribution model to accurately detect both known and unknown emerging SCAs at run-time. Our experimental results demonstrate that HybriDG achieves 100% detection accuracy for known attacks and 99.5% detection accuracy for unknown attacks which is significantly outperforming traditional ML algorithms, deep learning, and time series classification models by up to 80% for unknown and 8% known attack detection.
Han Wang 0020, Hossein Sayadi, Avesta Sasan, Setareh Rafatirad, Houman Homayoun
ICMLA4
2020 SCARF: Detecting Side-Channel Attacks at Real-time using Low-level Hardware Features
abstract
Side-Channel Attacks (SCAs) are powerful attacks compromising the security of modern computer systems by exploiting hardware vulnerabilities. Prior studies on detection of SCAs based on low-level microarchitectural features captured from processors' hardware performance counter (HPC) registers have considered collecting hardware events of both victim applications (cryptographic application, e.g. RSA, AES and etc.) and attack applications. However, in such techniques the attack HPCs data can be easily manipulated and/or corrupted resulting in misleading the SCA detection mechanism. Furthermore, the prior works have explored the suitability of a limited number of Machine Learning (ML) algorithms in detecting SCAs without examining the instance level false alarm rate that as we show in this work is a more important evaluation metric for SCA detection techniques. In response, in this paper, we propose SCARF, a machine learning-based real-time side-channel attack detection methodology using low-level hardware features. To this aim, we first only monitor the victim applications' behavior using the HPC features and analyze the captured low-level traces of the victim applications under no attack and attack conditions to avoid manipulation of attackers' HPCs. Next, a wide range of ML classifiers with customized HPC features are implemented to determine the most effective ML technique for detecting SCAs at real-time, while improving accuracy and reducing instance-level false alarm rate of ML-based SCA detectors. Lastly, the False Alarm Minimization (FAM) technique is proposed to further reduce the instance level false positive rate of the ML-based SCA detectors. The experimental results indicate that the SCARF methodology can obtain up to 100% attack detection accuracy with 0% instance level false alarm rate for detecting SCAs.
Han Wang 0020, Hossein Sayadi, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
IOLTS3
2020 Bidirectional Transformer based on online Text-based information to Implement Convolutional Neural Network Model For Secure Business Investment
abstract
Real estate investment decisions are critical for low-income people who have just one home as their life-time investment option. So during the COVID-19 pandemic, unemployment causes many homeowners with a low income to lose their homes because of two major factors: one, they could not pay their mortgages without a job, and second, their house could not be rented easily. Rent prediction in real-estate can guarantee the success of an investment. Online information from real estate websites plays a significant role in making a business decision to buy a home. This paper applies natural language processing models to introduce a new model for safe real estate investment based on online information. For the first time, we use a transfer learning model based on online information from various online resources to detect a profitable rental property. Bidirectional Encoder Representations from Transformers(BERT) are used to implement a semantic convolutional neural network model to predict real estate investment safety. This work introduces a new model for rent prediction based on the United States housing market. Our contribution is three-fold: (1) using natural language processing approach to use the semantics of online information on Airbnb, Zillow, Schools, Public transportation, and crime rate websites for rent prediction (2) We perform a comprehensive analysis of eager and lazy machine learning models as a traditional Machine learning models with our proposed new transfer learning model for rent prediction. (3) Creating a new public data set of semantic analysis for more than 5 million houses in the United States based on online information. This data set will be available for public research in natural language processing research for people analytic applications. This work introduces a new machine learning model to guarantee safe investment in the real estate market using a transfer learning approach based on online information.
Maryam Heidari, Setareh Rafatirad
ISTAS2
2019 XPPE: cross-platform performance estimation of hardware accelerators using machine learning
abstract
The increasing heterogeneity in the applications to be processed ceased ASICs to exist as the most efficient processing platform. Hybrid processing platforms such as CPU+FPGA are emerging as powerful processing platforms to support an efficient processing for a diverse range of applications. Hardware/Software co-design enabled designers to take advantage of these new hybrid platforms such as Zynq. However, dividing an application into two parts that one part runs on CPU and the other part is converted to a hardware accelerator implemented on FPGA, is making the platform selection difficult for the developers as there is a significant variation in the application's performance achieved on different platforms. Developers are required to fully implement the design on each platform to have an estimation of the performance. This process is tedious when the number of available platforms is large. To address such challenge, in this work we propose XPPE, a neural network based cross-platform performance estimation. XPPE utilizes the resource utilization of an application on a specific FPGA to estimate the performance on other FPGAs. The proposed estimation is performed for a wide range of applications and evaluated against a vast set of platforms. Moreover, XPPE enables developers to explore the design space without requiring to fully implement and map the application. Our evaluation results show that the correlation between the estimated speed up using XPPE and actual speedup of applications on a Hybrid platform over an ARM processor is more than 0.98.
Hosein Mohammadi Makrani, Hossein Sayadi, Tinoosh Mohsenin, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
ASP-DAC4
2019 Adversarial Attack on Microarchitectural Events based Malware Detectors
abstract
To overcome the performance overheads incurred by the traditional software-based malware detection techniques, Hardware-assisted Malware Detection (HMD) using machine learning (ML) classifiers has emerged as a panacea to detect malicious applications and secure the systems. To classify benign and malicious applications, HMD primarily relies on the generated low-level microarchitectural events captured through Hardware Performance Counters (HPCs). This work creates an adversarial attack on the HMD systems to tamper the security by introducing the perturbations in the HPC traces with the aid of an adversarial sample generator application. To craft the attack, we first deploy an adversarial sample predictor to predict the adversarial HPC pattern for a given application to be misclassified by the deployed ML classifier in the HMD. Further, as the attacker has no direct access to manipulate the HPCs generated during runtime, based on the output of the adversarial sample predictor, we devise an adversarial sample generator wrapped around a normal application to produce HPC patterns similar to the adversarial predictor HPC trace. As the crafted adversarial sample generator application does not have any malicious operations, it is not detectable with traditional signature-based malware detection solutions. With the proposed attack, malware detection accuracy has been reduced to 18.04% from 82.76%.
Sai Manoj Pudukotai Dinakarrao, Sairaj Amberkar, Sahil Bhat, Abhijitt Dhavlle, Hossein Sayadi, Avesta Sasan, Houman Homayoun, Setareh Rafatirad
DAC8
2019 Lightweight Node-level Malware Detection and Network-level Malware Confinement in IoT Networks
abstract
The sheer size of IoT networks being deployed today presents an "attack surface" and poses significant security risks at a scale never before encountered. In other words, a single device/node in a network that becomes infected with malware has the potential to spread malware across the network, eventually ceasing the network functionality. Simply detecting and quarantining the malware in IoT networks does not guarantee to prevent malware propagation. On the other hand, use of traditional control theory for malware confinement is not effective, as most of the existing works do not consider real-time malware control strategies that can be implemented using uncertain infection information of the nodes in the network or have the containment problem decoupled from network performance. In this work, we propose a two-pronged approach, where a runtime malware detector (HaRM) that employs Hardware Performance Counter (HPC) values to detect the malware and benign applications is devised. This information is fed during runtime to a stochastic model predictive controller to confine the malware propagation without hampering the network performance. With the proposed solution, a runtime malware detection accuracy of 92.21% with a runtime of 10ns is achieved, which is an order of magnitude faster than existing malware detection solutions. Synthesizing this output with the model predictive containment strategy lead to achieving an average network throughput of nearly 200% of that of IoT networks without any embedded defense.
Sai Manoj Pudukotai Dinakarrao, Hossein Sayadi, Hosein Mohammadi Makrani, Cameron Nowzari, Setareh Rafatirad, Houman Homayoun
DATE5
2019 2SMaRT: A Two-Stage Machine Learning-Based Approach for Run-Time Specialized Hardware-Assisted Malware Detection
abstract
Hardware-assisted Malware Detection (HMD) has emerged as a promising solution to improve the security of computer systems using Hardware Performance Counters (HPCs) information collected at run-time. While several recent studies proposed machine learning-based solutions to identify malware using HPCs, they rely on a large number of microarchitectural events to achieve high accuracy and detection rate. More importantly, they have largely overlooked complexity-effective prediction of malware classes at run-time. As we show in this work, the detection performance of malware classifiers is highly dependent on the number of available HPCs and varies significantly across classes of malware. The limited number of available HPCs in modern microprocessors that can be simultaneously captured makes run-time malware detection with high detection performance using existing solutions a challenging problem, as they require multiple runs of applications to collect a sufficient number of microarchitectural events. In response, in this paper, we first identify the most important HPCs for HMD using an effective feature reduction method. We then develop a specialized two-stage run-time HMD referred as 2SMaRT. 2SMaRT first classifies applications using a multiclass classification technique into either benign or one of the malware classes (Virus, Rootkit, Backdoor, and Trojan). In the second stage, to have a high detection performance, 2SMaRT deploys a machine learning model that works best for each class of malware. To realize an effective run-time solution that relies on only available HPCs, 2SMaRT is further customized using an ensemble learning technique to boost the performance of general malware detectors. The experimental results show that 2SMaRT using ensemble technique with just 4HPCs outperforms state-of-the-art classifiers with 8HPCs by up to 31.25% in terms of detection performance, on average across different classes of malware.
Hossein Sayadi, Hosein Mohammadi Makrani, Sai Manoj Pudukotai Dinakarrao, Tinoosh Mohsenin, Avesta Sasan, Setareh Rafatirad, Houman Homayoun
DATE6
2019 Pyramid: Machine Learning Framework to Estimate the Optimal Timing and Resource Usage of a High-Level Synthesis Design
abstract
The emergence of High-Level Synthesis (HLS) tools shifted the paradigm of hardware design by making the process of mapping high-level programming languages to hardware design such as C to VHDL/Verilog feasible. HLS tools offer a plethora of techniques to optimize designs for both area and performance, but resource usage and timing reports of HLS tools mostly deviate from the post-implementation results. In addition, to evaluate a hardware design performance, it is critical to determine the maximum achievable clock frequency. Obtaining such information using static timing analysis provided by CAD tools is difficult, due to the multitude of tool options. Moreover, a binary search to find the maximum frequency is tedious, time-consuming, and often does not obtain the optimal result. To address these challenges, we propose a framework, called Pyramid, that uses machine learning to accurately estimate the optimal performance and resource utilization of an HLS design. For this purpose, we first create a database of C-to- FPGA results from a diverse set of benchmarks. To find the achievable maximum clock frequency, we use Minerva, which is an automated hardware optimization tool. Minerva determines the close-to-optimal settings of tools, using static timing analysis and a heuristic algorithm, and targets either optimal throughput or throughput-to-area. Pyramid uses the database to train an ensemble machine learning model to map the HLS-reported features to the results of Minerva. To this end, Pyramid recalibrates the results of HLS to bridge the accuracy gap, and enable developers to estimate the throughput or throughputto- area of hardware design with more than 95% accuracy and alleviates the need to perform actual implementation for estimation.
Hosein Mohammadi Makrani, Farnoud Farahmand, Hossein Sayadi, Sara Bondi, Sai Manoj Pudukotai Dinakarrao, Houman Homayoun, Setareh Rafatirad
FPL7
2019 On Custom LUT-based Obfuscation
abstract
Logic obfuscation yields hardware security against various threats, such as Intellectual Property (IP) piracy and reverse engineering. Evolving Boolean satisfiability (SAT) attacks have challenged the hardware security assurance rendered by various obfuscation methods. Recent works have centered on using re-configurable components such as Look-Up-Tables (LUTs) to enhance resiliency against attacks. Resiliency against SAT-attack is guaranteed when the size of LUT (number of inputs) is large. However, this incurs significant power, area and performance overheads. To address this challenge, this work proposes logic encryption based on customized LUT to make this practical. We propose two variants of the customized LUT based obfuscation: LUT+MUX based obfuscation, securing the design through routing obfuscation by MUX(multiplexer) and logic obfuscation of LUTs; and LUT+LUT based obfuscation, benefiting from LUT based obfuscation reinforced with additional logic/routing obfuscation. We evaluate the hardware security and overheads of the proposed two variants of customized LUT-based obfuscation on various benchmarks. Proposedcustomized LUT-based obfuscation breaks the security, power, and area trade-offs. The proposed solution is shown to be robust against SAT-attacks and power analysis-based side-channel attacks with8×reduced area and 3×reduced power on an average compared tostate-of-the-art LUT-based obfuscation.
Gaurav Kolhe, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Hamid Mahmoodi, Avesta Sasan, Houman Homayoun
ACM Great Lakes Symposium on VLSI3
2019 Mitigating the Performance and Quality of Parallelized Compressive Sensing Reconstruction Using Image Stitching
abstract
Orthogonal Matching Pursuit is an iterative greedy algorithm used to find a sparse approximation for high-dimensional signals. The algorithm is most popularly used in Compressive Sensing, which allows for the reconstruction of sparse signals at rates lower than the Shannon-Nyquist frequency, which has traditionally been used in a number of applications such as MRI and computer vision and is increasingly finding its way into Big Data and data center analytics. OMP traditionally suffers from being computationally intensive and time-consuming, this is particularly a problem in the area of Big Data where the demand for computational resources continues to grow. In this paper, the data-level parallelization of OMP through blocking is examined. Traditionally blocking has been used to ac- celerate the performance of OMP reconstruction for big data image analytics. However, as we show in this work, blocking, particularly in the form of vectorizing, introduces significant error in terms of PSNR and SSIM index in the reconstruction quality. In response, we deploy the concept of stitching to recover the lost accuracy. We further examine the influence of the level of blocking and amount of stitching (overlap between each block) with regard to recon- struction time and reconstructed image quality. While stitching boosts up the image reconstruction accuracy significantly, the ob- ject detection count results show anywhere from 11.84% to 140.54% improvement, depending on the cases being compared, it introduces significant overhead with regard to reconstruction time. To address the overhead, we deploy hardware accelerated base solutions. Given the emergence of hardware accelerators in data centers and for big data analytics in form of FPGAs, our solution effectively utilizes this resource to enhance the performance overhead of stitching by 25%. We show the minimum block size required for an FPGA speed-up.
Mahmoud Namazi, Hosein Mohammadi Makrani, Zhi Tian, Setareh Rafatirad, Mohamad Hosein Akbari, Avesta Sasan, Houman Homayoun
ACM Great Lakes Symposium on VLSI4
2019 Security and Complexity Analysis of LUT-based Obfuscation: From Blueprint to Reality
abstract
Recent obfuscation schemes have leveraged reconfigurable logics to alleviate various hardware security threats. However, existing reconfigurable logic-based obfuscation schemes focus on specific design factors such as gate replacement strategy or an optimization metric such as SAT-hardness. Despite meeting the focused metrics such as security, the obfuscation also incurs overheads, which are not well analyzed in the existing works. In this work, we provide a comprehensive analysis on reconfigurable logic obfuscation schemes i.e., LUT-based obfuscation by investigating 3-key design factors such as (1) LUT size, (2) number of LUTs, and (3) replacement strategy as they have a considerable impact on design criteria, i.e., Power-Performance-Area (PPA) and Security (PPA/S). Our results show that among the studied parameters the size of LUT has the most prominent impact on improving the resiliency of LUT-based obfuscation against the SAT and removal attacks. However, using large size LUTs incur significant PPA overheads, making such solutions unfeasible and unpractical. To address this challenge, this work proposes a pragmatic solution based on a customized LUT, where the security provided by each LUT is superior to that of traditional LUT-based obfuscation. The proposed solution primarily benefits from LUT-based obfuscation reinforced with additional logic/routing obfuscation that is implemented using small 2-input LUTs. We evaluate the hardware security and overhead of the proposed customized LUT-based obfuscation on various benchmarks to prove that the customized LUT-based obfuscation breaks the PPA tradeoffs while exhibiting robustness against the SAT and removal attacks. The customized LUT-based obfuscation comes with 8× reduced area and 2× reduced power on an average compared to state-of-the-art LUT-based obfuscation without compromising security.
Gaurav Kolhe, Hadi Mardani Kamali, Miklesh Naicker, Tyler David Sheaves, Hamid Mahmoodi, Sai Manoj Pudukotai Dinakarrao, Houman Homayoun, Setareh Rafatirad, Avesta Sasan
ICCAD8
2019 Deep Multi-attributed Graph Translation with Node-Edge Co-Evolution
abstract
Generalized from image and language translation, graph translation aims to generate a graph in the target domain by conditioning an input graph in the source domain. This promising topic has attracted fast-increasing attention recently. Existing works are limited to either merely predicting the node attributes of graphs with fixed topology or predicting only the graph topology without considering node attributes, but cannot simultaneously predict both of them, due to substantial challenges: 1) difficulty in characterizing the interactive, iterative, and asynchronous translation process of both nodes and edges and 2) difficulty in discovering and maintaining the inherent consistency between the node and edge in predicted graphs. These challenges prevent a generic, end-to-end framework for joint node and edge attributes prediction, which is a need for real-world applications such as malware confinement in IoT networks and structural-to-functional network translation. These real-world applications highly depend on hand-crafting and ad-hoc heuristic models, but cannot sufficiently utilize massive historical data. In this paper, we termed this generic problem "multi-attributed graph translation" and developed a novel framework integrating both node and edge translations seamlessly. The novel edge translation path is generic which is proven to be a generalization of the existing topology translation models. Then, a spectral graph regularization based on our non-parametric graph Laplacian is proposed to learn and maintain the consistency of the predicted nodes and edges. Finally, extensive experiments on both synthetic and real-world application data demonstrated the effectiveness of the proposed method.
Xiaojie Guo 0002, Liang Zhao 0002, Cameron Nowzari, Setareh Rafatirad, Houman Homayoun, Sai Manoj Pudukotai Dinakarrao
ICDM4
2019 RNN-Based Classifier to Detect Stealthy Malware using Localized Features and Complex Symbolic Sequence
abstract
Malware detection and classification has enticed a lot of researchers in the past decades. Several mechanisms based on machine learning (ML), computer vision and deep learning have been deployed to this task and have achieved considerable results. However, advanced malware (stealthy malware) generated using various obfuscation techniques like code relocation, code transposition, polymorphism and mutation thwart the detection. In this paper, we propose a two-pronged technique which can efficiently detect both traditional and stealthy malware. Firstly, we extract the microarchitectural traces procured while executing the application, which are fed to the traditional ML classifiers to identify malware spawned as separate thread. In parallel, for an efficient stealthy malware detection, we instigate an automated localized feature extraction technique that will be used as an input to recurrent neural networks (RNNs) for classification. We have tested the proposed mechanism rigorously on stealthy malware created using code relocation obfuscation technique. With the proposed two-pronged approach, an accuracy of 94%, precision of 93%, recall score of 96% and F-1 score of 94% is achieved. Furthermore, the proposed technique attains up to 11% higher on average detection accuracy and precision, along with 24% higher on average recall and F-1 score as compared to the CNN-based sequence classification and hidden Markov model (HMM) based approaches in detecting stealthy malware.
Sanket Shukla, Gaurav Kolhe, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad
ICMLA4
2019 A+ Tuning: Architecture+Application Auto-Tuning for In-Memory Data-Processing Frameworks
abstract
Processing big data eventually leads to an upsurge in datacenters' power consumption, which is one of the pivotal concerns to be addressed. Many of the existing works focus on optimizing either power or performance, which is not the best parameter to consider for achieving high energy efficiency with low operational costs. Furthermore, the existing works require profiling of big data applications exhaustively and only consider tuning of either architectural or software parameters, often leading to sub-optimal settings. To cope up with the above-mentioned drawbacks of the existing works, we propose a system, A+ Tuning (Architecture + Application Auto-tuning) which enables us to determine a close to optimal settings by simultaneously optimizing for Energy Delay Product(EDP), representing energy efficiency. The proposed A+ Tuning involves a) profile the incoming unknown applications to different types (compute-bound, memory-bound and etc.) based on known applications classification result; b) co-locate the applications, and c) employs a machine learning-based model to determine the optimal settings and tune from both architectural and application settings for the co-located applications. By applying the proposed A+Tuning system, datacenters achieve up to 4×EDP improvement compared to fairshare methodology and 2.5×compared with recent works such as BestConfig.
Han Wang 0020, Setareh Rafatirad, Houman Homayoun
ICPADS2
2019 ECoST: Energy-Efficient Co-Locating and Self-Tuning MapReduce Applications
abstract
Datacenters provide high performance and flexibility for users and cost efficiency for operators. Hyperscale datacenters are harnessing massively scalable computer resources for large-scale data analysis. However, cloud/datacenter infrastructure does not scale as fast as the input data volume and computational requirements of big data and analytics technologies. Thus, more applications need to share CPU at the node level that could have large impact on performance and operational cost. To address this challenge, in this paper we show that, concurrently fine-tune parameters at the application, microarchitecture, and system levels are creating opportunities to co-locate applications at the node level and improve energy-efficiency of the server while maintaining performance. Co-locating and self-tuning of unknown applications are challenging problems, especially when co-locating multiple big data applications concurrently with many tuning knobs, potentially requiring exhaustive brute-force search to find the right settings. This research challenge upsurges an imminent need to develop a technique that co-locates applications at a node level and predict the optimal system, architecture and application level configure parameters to achieve the maximum energy efficiency. It promotes the scale-down of computational nodes by presenting the Energy-Efficient Co-Locating and Self-Tuning (ECoST) technique for data intensive applications. ECoST proof of concept was successfully tested on MapReduce platform. ECoST can also be deployed on other data-intensive frameworks where there are several parameters for power and performance tuning optimizations. ECoST collects run-time hardware performance counter data and implements various machine learning models from as simple as a lookup table or decision tree based to as complex as neural network based to predict the energy-efficiency of co-located applications. Experimental data show energy efficiency is achieved within 4% of the upper bound results when co-locating multiple applications at a node level. ECoST is also scalable, being within 8% of upper bound on an 8-node server.
Maria Malik, Hassan Ghasemzadeh 0001, Tinoosh Mohsenin, Rosario Cammarota, Liang Zhao 0002, Avesta Sasan, Houman Homayoun, Setareh Rafatirad
ICPP8
2019 Stealthy Malware Detection using RNN-Based Automated Localized Feature Extraction and Classifier
abstract
Malware analysis, detection and classification has allured a lot of researchers in the past few years. Numerous methods based on machine learning (ML), computer vision and deep learning have been applied to this task and have accomplished some pragmatic results. One of the basic assumption of these works is that malware is spawned as a separate thread and the distinguishing features can be extracted in a "clean" manner irrespective of the malware obfuscation deployed. However, this assumption does not hold true for the advanced malware obfuscation techniques such as code relocation, mutation and polymorphism. Stealthy malware is a malware created by embedding the malware in a benign application through advanced obfuscation strategies to thwart the detection. To perform efficient malware detection for traditional and stealthy malware alike, we propose a two-pronged approach. Firstly, we extract the microarchitectural traces obtained while executing the application, which are fed to the traditional ML classifiers to detect malware spawned as separate thread. In parallel, for an efficient stealthy malware detection, we introduce an automated localized feature extraction technique that will be further processed using the recurrent neural networks (RNNs) for classification. To perform this, we translate the application binaries into images and further convert it into sequences and extract local features for stealthy malware detection. With the proposed two-pronged approach, an accuracy of 94% and nearly 90% is achieved in detecting normal and stealthy malware created through code relocation obfuscation technique. Furthermore, the proposed approach achieves up to 11% higher detection accuracy compared to the CNN-based sequence classification and hidden Markov model (HMM) based approaches in detecting stealthy malware.
Sanket Shukla, Gaurav Kolhe, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad
ICTAI4
2019 Big vs little core for energy-efficient Hadoop computing
Maria Malik, Katayoun Neshatpour, Setareh Rafatirad, Rajiv V. Joshi, Tinoosh Mohsenin, Hassan Ghasemzadeh 0001, Houman Homayoun
J. Parallel Distributed Comput.3
2018 Compressive Sensing on Storage Data: An Effective Solution to Alleviate I/0 Bottleneck in Data- Intensive Workloads
abstract
The gap between computation speed and I/O access on modern computing systems imposes processing limitations in data-intensive applications. Employing high-end memory has proven not to enhance the performance for I/O bound applications, given the low utilization of memory bandwidth in such applications, as highlighted in recent studies. Despite several solutions to improve the performance of storage, none of them is able to shift the bottleneck from the I/O access to the memory subsystem for I/O bound applications. In this paper, we show that in the case of data-intensive multimedia applications, by using Compressive Sensing (CS), a lossy data compression method, the bottleneck is lifted from the storage, increasing the bandwidth utilization of the memory to gain further performance improvement from a high-end memory. The reconstruction of compressed data is however time and memory consuming. To address this challenge, we employ and compare the hardware and software acceleration of Orthogonal Matching Pursuit (OMP), a greedy algorithm, which solves the problem by choosing the most significant variable to reduce the least square error. Our implementation results show that CS increases memory bandwidth utilization by 1.4x and using high bandwidth memory results in 24% performance improvement. Overall, the proposed solution of CS of storage data with FPGA accelerator achieves up to 45% speedup in an end-to-end implementation by only 4.6% accuracy degradation.
Hosein Mohammadi Makrani, Hossein Sayadi, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Houman Homayoun
ASAP4
2018 Advances and throwbacks in hardware-assisted security: special session
Ferdinand Brasser, Lucas Davi, Abhijitt Dhavlle, Tommaso Frassetto, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Ahmad-Reza Sadeghi, Avesta Sasan, Hossein Sayadi, Shaza Zeitouni, Houman Homayoun
CASES6
2018 Main-Memory Requirements of Big Data Applications on Commodity Server Platform
abstract
The emergence of big data frameworks requires computational and memory resources that can naturally scale to manage massive amounts of diverse data. It is currently unclear whether big data frameworks such as Hadoop, Spark, and MPI will require high bandwidth and large capacity memory to cope with this change. The primary purpose of this study is to answer this question through empirical analysis of different memory configurations available for commodity server and to assess the impact of these configurations on the performance Hadoop and Spark frameworks, and MPI based applications. Our results show that neither DRAM capacity, frequency, nor the number of channels play a critical role on the performance of all studied Hadoop as well as most studied Spark applications. However, our results reveal that iterative tasks (e.g. machine learning) in Spark and MPI are benefiting from a high bandwidth and large capacity memory.
Hosein Mohammadi Makrani, Setareh Rafatirad, Amir Houmansadr, Houman Homayoun
CCGrid2
2018 Comprehensive assessment of run-time hardware-supported malware detection using general and ensemble learning
abstract
Recent studies have demonstrated the effectiveness of Hardware Performance Counters (HPCs) for detecting pattern of malicious applications. Hardware-supported detectors utilize Machine Learning (ML) classifiers for malware detection by analyzing a large number of HPC features, more than the very limited number of HPC registers available in modern microprocessors. Obtaining more HPCs requires running the application (malware or benign) more than once to collect the required data, which in turn makes the solution less practical for run-time detection of malware. In response to this challenge, in this work, we first identify the critical HPC features required for malware detection. Next, we explore the use of various ML techniques to classify benign and malware applications using the selected HPCs at run-time. Further, we investigate the effectiveness of ensemble learning in improving the performance of ML classifiers. For this purpose, we apply AdaBoost on all general ML classifiers. We thoroughly compare the general and ensemble ML classifiers in terms of accuracy, robustness, performance, and hardware overhead. The experimental results indicate that ensemble learning enhances the performance of malware detection for rule-based and tree-based algorithms up to 13%. However, it diminishes the performance of neural network and Bayesian network-based detectors by 6% and 4%, respectively.
Hossein Sayadi, Sai Manoj Pudukotai Dinakarrao, Amir Houmansadr, Setareh Rafatirad, Houman Homayoun
CF4
2018 Energy-aware and Machine Learning-based Resource Provisioning of In-Memory Analytics on Cloud
abstract
In this work, we propose a proactive online resource provisioning methodology that addresses the challenge of resource provisioning for IMC workloads in heterogeneous cloud platforms consist of diverse types of servers. As cloud platforms provide a wide range of server configuration choices [4], and the applications' performance and power consumption changes at run-time [3] and depends on the chosen configuration, resource provisioning in cloud platforms is a challenging optimization problem with a large search space to navigate. Our methodology proactively assigns a suitable hardware configuration to IMC program for energy-efficiency (EDP) optimization at run-time before any significant change occurs in application's behavior. This helps to save energy without sacrificing performance [2, 7]. We address these challenges by first characterizing diverse types of IMC workloads across different types of server architectures. The characterization aids to accurately capture applications' behavior [1] and train machine learning models [5, 6]. We use time series neural network to predict the next phase of an application. Our approach then uses artificial neural networks to estimate the performance and power consumption of predicted phase of application on various server configurations. Further, we use the genetic algorithm to distinguish close-to-optimal configuration to minimize EDP. Compared to Oracle scheduler, our methodology achieves 93% accuracy to allocate the right resource for each phase of the program. Our methodology improves the performance by 21% and the EDP by 40% on average, compared to the default scheduler.
Hosein Mohammadi Makrani, Hossein Sayadi, Devang Motwani, Han Wang 0020, Setareh Rafatirad, Houman Homayoun
SoCC5
2018 Ensemble learning for effective run-time hardware-based malware detection: a comprehensive analysis and classification
abstract
Malware detection at the hardware level has emerged recently as a promising solution to improve the security of computing systems. Hardware-based malware detectors take advantage of Machine Learning (ML) classifiers to detect pattern of malicious applications at run-time. These ML classifiers are trained using low-level features such as processor Hardware Performance Counters (HPCs) data which are captured at run-time to appropriately represent the application behaviour. Recent studies show the potential of standard ML-based classifiers for detecting malware using analysis of large number of microarchitectural events, more than the very limited number of HPC registers available in today's microprocessors which varies from 2 to 8. This results in executing the application more than once to collect the required data, which in turn makes the solution less practical for effective run-time malware detection. Our results show a clear trade-off between the performance of standard ML classifiers and the number and diversity of HPCs available in modern microprocessors. This paper proposes a machine learning-based solution to break this trade-off to realize effective run-time detection of malware. We propose ensemble learning techniques to improve the performance of the hardware-based malware detectors despite using a very small number of microarchitectural events that are captured at run-time by existing HPCs, eliminating the need to run an application several times. For this purpose, eight robust machine learning models and two well-known ensemble learning classifiers applied on all studied ML models (sixteen in total) are implemented for malware detection and precisely compared and characterized in terms of detection accuracy, robustness, performance (accuracy×robustness), and hardware overheads. The experimental results show that the proposed ensemble learning-based malware detection with just 2 HPCs using ensemble technique outperforms standard classifiers with 8 HPCs by up to 17%. In addition, it can match the robustness and performance of standard ML-based detectors with 16 HPCs while using only 4 HPCs allowing effective run-time detection of malware.
Hossein Sayadi, Nisarg Patel, Sai Manoj Pudukotai Dinakarrao, Avesta Sasan, Setareh Rafatirad, Houman Homayoun
DAC5
2018 Design Space Exploration for Hardware Acceleration of Machine Learning Applications in MapReduce
abstract
Emerging big data applications heavily rely on machine learning algorithms which are computationally intensive. To meet computational requirements, and power and scalability challenges, FPGA based Hardware accelerators have found their way in data centers and cloud infrastructures. Recent efforts on HW acceleration of big data mainly attempt to accelerate a particular application and deploy it on a specific architecture that fits well its performance and power requirements. Given the diversity of architectures and ML applications, the important research question is which architecture is better suited to meet the performance, power and energy-efficiency requirements of a diverse range of ML-based analytics applications. In this work, we answer this question by investigating how the type of FPGA (low-end vs. high-end), and its integration with the CPU (on-chip vs. off-chip) along with the choice of CPU (high performance big vs. low power little servers) affects the speedup yield and power reduction in a CPU+FPGA architecture for machine learning applications implemented in MapReduce. We show that among the three architectural parameters, the type of CPU is the most dominant factor in determining the execution time and power in a CPU+FPGA architecture for MapReduce applications. The integration technology and FPGA type comes next, with the power and performance least sensitive to the FPGA type.
Katayoun Neshatpour, Hosein Mohammadi Makrani, Avesta Sasan, Hassan Ghasemzadeh 0001, Setareh Rafatirad, Houman Homayoun
FCCM5
2018 Efficient utilization of adversarial training towards robust machine learners and its analysis
abstract
Advancements in machine learning led to its adoption into numerous applications ranging from computer vision to security. Despite the achieved advancements in the machine learning, the vulnerabilities in those techniques are as well exploited. Adversarial samples are the samples generated by adding crafted perturbations to the normal input samples. An overview of different techniques to generate adversarial samples, defense to make classifiers robust is presented in this work. Furthermore, the adversarial learning and its effective utilization to enhance the robustness and the required constraints are experimentally provided, such as up to 97.65% accuracy even against CW attack. Though adversarial learning's effectiveness is enhanced, still it is shown in this work that it can be further exploited for vulnerabilities.
Sai Manoj Pudukotai Dinakarrao, Sairaj Amberkar, Setareh Rafatirad, Houman Homayoun
ICCAD3
2018 Energy-efficient acceleration of MapReduce applications using FPGAs
Katayoun Neshatpour, Maria Malik, Avesta Sasan, Setareh Rafatirad, Tinoosh Mohsenin, Hassan Ghasemzadeh 0001, Houman Homayoun
J. Parallel Distributed Comput.4
2018 Optimal Allocation of Computation and Communication in an IoT Network
abstract
Internet of things (IoT) is being developed for a wide range of applications from home automation and personal fitness to smart cities. With the extensive growth in adaptation of IoT devices comes the uncoordinated and substandard designs aimed at promptly making products available to the end consumer. This substandard approach restricts the growth of IoT in the near future and necessitates that studies understand requirements for an efficient design. A particular area where IoT applications have grown significantly is surveillance and monitoring. Applications of IoT in this domain are relying on distributed sensors, each equipped with a battery, capable of collecting images, processing images, and communicating the raw or processed data to the nearest node until it reaches the base station for decision making. In such an IoT network where processing can be distributed over the network, the important research question is how much of data each node should process and how much it should communicate for a given objective. This work answers this question and provides a deeper understanding of energy and delay tradeoffs in an IoT network with three different target metrics.
Abhimanyu Chopra, Hakan Aydin, Setareh Rafatirad, Houman Homayoun
ACM Trans. Design Autom. Electr. Syst.3
2018 Programmable Gates Using Hybrid CMOS-STT Design to Prevent IC Reverse Engineering
abstract
This article presents a rigorous step towards design-for-assurance by introducing a new class of logically reconfigurable design resilient to design reverse engineering. Based on the non-volatile spin transfer torque (STT) magnetic technology, we introduce a basic set of non-volatile reconfigurable Look-Up-Table (LUT) logic components (NV-STT-based LUTs). An STT-based LUT with a significantly different set of characteristics compared to CMOS provides new opportunities to enhance design security yet makes it challenging to remain highly competitive with custom CMOS or even SRAM-based LUT in terms of power, performance, and area. To address these challenges, we propose several algorithms to select and replace custom CMOS gates with reconfigurable STT-based LUTs during design implementation such that the functionality of STT-based components and therefore the entire design cannot be determined in any manageable time, rendering any design reverse engineering attack ineffective. Our study, conducted on a large number of standard circuit benchmarks, concludes significant resiliency of hybrid STT-CMOS circuits against various types of attacks. Furthermore, the selection algorithms on average have a small impact on the performance of the circuit. We also tested these techniques against satisfiability attacks developed recently and show that these techniques also render more advanced reverse-engineering techniques computationally infeasible.
Theodore Winograd, Gaurav Shenoy, Hassan Salmani, Hamid Mahmoodi, Setareh Rafatirad, Houman Homayoun
ACM Trans. Design Autom. Electr. Syst.5
2016 Big biomedical image processing hardware acceleration: A case study for K-means and image filtering
abstract
Most hospitals today are dealing with the big data problem, as they generate and store petabytes of patient records most of which in form of medical imaging, such as pathological images, CT scans and X-rays in their datacenters. Analyzing such large amounts of biomedical imaging data to enable discovery and guide physicians in personalized care is becoming an important focus of data mining and machine learning algorithms developed for biomedical Informatics (BMI). Algorithms that are developed for BMI heavily rely on complex and computationally intensive machine learning and data mining methods to learn from large data. The high processing demand of big biomedical imaging data has given rise to their implementation in high-end server platforms running software ecosystems that are optimized for dealing with large amount of data including Apache Hadoop and Apache Spark. However, efficient processing of such large amount of imaging data running computational intensive learning methods is becoming a challenging problem using state-of-the-art high performance computing server architectures. To address this challenge, in this paper, we introduce a scalable and efficient hardware acceleration method using low cost commodity FPGAs that is interfaced with a server architecture through a high speed interface. In this work we present a full end-to-end implementation of big data image processing and machine learning applications in a heterogeneous CPU+FPGA architecture. We develop the MapReduce implementation of K-means and Laplacian Filtering in Hadoop Streaming environment that allows developing mapper functions in non-Java based languages suited for interfacing with FPGA-based hardware accelerating environment. We accelerate the mapper functions through hardware+software (HW+SW) co-design. We do a full implementation of the HW+SW mappers on the Zynq FPGA platform. The results show promising kernel speedup of up to 27× for large image data sets. This translate to 7.8× and 1.8× speedup in an end-to-end Hadoop MapReduce implementation of K-mean s and Laplacian Filtering algorithm, respectively.
Katayoun Neshatpour, Arezou Koohi, Farnoud Farahmand, Rajiv V. Joshi, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
ISCAS5
2016 Characterizing Hadoop applications on microservers for performance and energy efficiency optimizations
abstract
The traditional low-power embedded processors such as Atom and ARM are entering the high-performance server market. At the same time, as the size of data grows, emerging Big Data applications require more and more server computational power that yields challenges to process data energy-efficiently using current high performance server architectures. Furthermore, physical design constraints, such as power and density have become the dominant limiting factor for scaling out servers. Numerous big data applications rely on using the Hadoop MapReduce framework to perform their analysis on large-scale datasets. Since Hadoop configuration parameters as well as architecture parameters directly affect the MapReduce job performance and energy-efficiency, system and architecture level parameters tuning is vital to maximize the energy efficiency. In this work, through methodical investigation of performance and power measurements, we demonstrate how the interplay among various Hadoop configurations and system and architecture level parameters affect the performance and energy-efficiency across various Hadoop applications.
Maria Malik, Avesta Sasan, Rajiv V. Joshi, Setareh Rafatirad, Houman Homayoun
ISPASS4
2015 System and architecture level characterization of big data applications on big and little core server architectures
abstract
Emerging Big Data applications require a significant amount of server computational power. Big data analytics applications rely heavily on specific deep machine learning and data mining algorithms, and exhibit high computational intensity, memory intensity, I/O intensity and control intensity. Big data applications require computing resources that can efficiently scale to manage massive amounts of diverse data. However, the rapid growth in the data yields challenges to process data efficiently using current server architectures such as big Xeon cores. Furthermore, physical design constraints, such as power and density, have become the dominant limiting factor for scaling out servers. Therefore recent work advocates the use of low-power embedded cores in servers such as little Atom to address these challenges. In this work, through methodical investigation of power and performance measurements, and comprehensive system level and micro-architectural analysis, we characterize emerging big data applications on big Xeon and little Atom-based server architecture. The characterization results across a wide range of real-world big data applications and various software stacks demonstrate how the choice of big vs little core-based server for energy-efficiency is significantly influenced by the size of data, performance constraints, and presence of accelerator. Furthermore, the microarchitecture-level analysis highlights where improvement is needed in big and little cores microarchitecture.
Maria Malik, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
IEEE BigData2
2011 A comprehensive study of visual event computing
Wei Qi Yan 0001, Declan F. Kieran, Setareh Rafatirad, Ramesh Jain 0001
Multim. Tools Appl.3
2009 MEDIALIFE: from images to a life chronicle
abstract
demonstration Share on MEDIALIFE: from images to a life chronicle Authors: Amarnath Gupta University of California San Diego, La Jolla, CA, USA University of California San Diego, La Jolla, CA, USAView Profile , Setareh Rafatirad University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile , Mingyan Gao University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile , Ramesh Jain University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile Authors Info & Claims SIGMOD '09: Proceedings of the 2009 ACM SIGMOD International Conference on Management of dataJune 2009 Pages 1119–1122https://doi.org/10.1145/1559845.1559998Published:29 June 2009Publication History 4citation280DownloadsMetricsTotal Citations4Total Downloads280Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Amarnath Gupta, Setareh Rafatirad, Mingyan Gao, Ramesh Jain 0001
SIGMOD Conference2