VLDB 2026 Research / reviewers in the wild / expert
Najmeh Nazari
dblp:251/4305 · also Najmeh Nazari Bavarsad
· DBLP profile ↗
21ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0003-3094-9439ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 9 first-author · 12 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bit of a Close Talker: A Practical Guide to Serverless Cloud Co-Location Attacks
Najmeh Nazari, Behnam Omidi, Setareh Rafatirad, Khaled N. Khasawneh, Houman Homayoun, Chongzhou Fang |
NDSS | 2 |
| 2025 | Transformers for Secure Hardware Systems: Applications, Challenges, and Outlook
Banafsheh S. Latibari, Najmeh Nazari, Avesta Sasan, Houman Homayoun, Pratik Satam, Soheil Salehi, Hossein Sayadi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical ImagingabstractVision Transformers (ViTs) have emerged as powerful architectures in medical image analysis, excelling in tasks such as disease detection, segmentation, and classification. However, their reliance on large, attention-driven models makes them vulnerable to hardware-level attacks. In this paper, we propose a novel threat model referred to as Med-Hammer that combines the Rowhammer hardware fault injection with neural Trojan attacks to compromise the integrity of ViT-based medical imaging systems. Specifically, we demonstrate how malicious bit flips induced via Rowhammer can trigger implanted neural Trojans, leading to targeted misclassification or suppression of critical diagnoses (e.g., tumors or lesions) in medical scans. Through extensive experiments on benchmark medical imaging datasets such as ISIC, Braib Tumor, and MedMNIST, we show that such attacks can remain stealthy while achieving high attack success rates about 82.51% and 92.56% in MobileViT and SwinTransformer, respectively. We further investigate how architectural properties, such as model sparsity, attention weight distribution, and number of features of the layer, impact attack effectiveness. Our findings highlight a critical and underexplored intersection between hardware-level faults and deep learning security in healthcare applications, underscoring the urgent need for robust defenses spanning both model architectures and underlying hardware platforms. Banafsheh S. Latibari, Najmeh Nazari, Hossein Sayadi, Houman Homayoun, Abhijit Mahalanobis |
ICCD | 2 |
| 2025 | FaRAccel: FPGA-Accelerated Defense Architecture for Efficient Bit-Flip Attack Resilience in Transformer ModelsabstractForget and Rewire (FaR) methodology has demonstrated strong resilience against Bit-Flip Attacks (BFAs) on Transformer-based models by obfuscating critical parameters through dynamic rewiring of linear layers. However, the application of FaR introduces non-negligible performance and memory overheads, primarily due to the runtime modification of activation pathways and the lack of hardware-level optimization. To overcome these limitations, we propose FaRAccel, a novel hardware accelerator architecture implemented on FPGA, specifically designed to offload and optimize FaR operations. FaRAccel integrates reconfigurable logic for dynamic activation rerouting, and lightweight storage of rewiring configurations, enabling low-latency inference with minimal energy overhead. We evaluate FaRAccel across a suite of Transformer models and demonstrate substantial reductions in FaR inference latency and improvement in energy efficiency, while maintaining the robustness gains of the original FaR methodology. To the best of our knowledge, this is the first hardware-accelerated defense against BFAs in Transformers, effectively bridging the gap between algorithmic resilience and efficient deployment on real-world AI platforms. Najmeh Nazari, Banafsheh S. Latibari, Elahe Hosseini, Fatemeh Movafagh, Chongzhou Fang, Hosein Mohammadi Makrani, Kevin Immanuel Gubbi, Abhijit Mahalanobis, Setareh Rafatirad, Hossein Sayadi, Houman Homayoun |
ICCD | 1 |
| 2025 | Extended Operational Life for Wearable Health Devices: A Hybrid TinyML and Server-Side ML ApproachabstractWearable devices equipped with sensors like ECG, PPG, and heart rate monitors are pivotal in health monitoring, yet they face significant challenges due to limited battery life, particularly when using complex machine learning (ML) models for continuous monitoring. This paper explores an innovative approach to enhance energy efficiency in wearable health monitoring applications by integrating TinyML with advanced ML models. We propose using energy-efficient TinyML models for initial binary classification to detect abnormalities in ECG signals. Upon detecting an anomaly, the data is transmitted to a server where more complex ML models, perform detailed multi-label classification to identify specific conditions. Our methodology leverages event-driven software frameworks and adaptive data collection strategies to optimize power consumption, extending the operational life of wearable devices without sacrificing accuracy. The proposed system is evaluated on several metrics, including accuracy, F1-score, energy consumption, and model latency. Our findings demonstrate that the integration of TinyML can significantly extend battery life while maintaining high levels of accuracy and reliability in health monitoring, presenting a promising solution for long-term wearable device deployment. Najmeh Nazari, Vedant Patel, Chongzhou Fang, Setareh Rafatirad, Houman Homayoun |
ISCAS | 1 |
| 2025 | Large Language Models for Opioid-Induced Respiratory Depression Prediction in Hospitalized Patients: A Retrospective StudyabstractOpioid-induced adverse events pose significant risks to hospitalized patients. However, there is a limited understanding of which patients on general care floors are at risk for Opioid-Induced Respiratory Depression (OIRD). This study aims to bridge that knowledge gap by utilizing the advancements of AI to interpret Electronic Medical Records (EMRs). Recently, Large Language Models (LLMs) have gained attention for their exceptional capabilities in understanding human language, which makes them crucial for AI systems in healthcare that focus on clinical narratives. In this study, we extracted 2,663 hospitalized adult patient records from UC Davis Medical Center archives between January 2010 and April 2020 to identify patients at high risk of OIRD. For this purpose, we employed clinical language models (BioBERT, ClinicalBERT, and GatorTron) and fine-tuned them on the OIRD dataset. Additionally, we leveraged the capabilities of GPT-4, a state-of-the-art LLM, to select the most informative risk factors and enhance the accuracy of the predictive models. Elahe Hosseini, Abhinav Srinivas, Najmeh Nazari, Charity Hale, Setareh Rafatirad, Houman Homayoun |
ACM Trans. Comput. Heal. | 3 |
| 2025 | MLB-MAC: Multi-Level Binary MAC Array for Energy Efficient ML AcceleratorsabstractQuantization is a critical compression technique for optimizing deep neural networks (DNNs) on resource-constrained embedded devices. Efficient hardware utilization hinges on effective number representation. Integer representation, a widely adopted method, uses scaling factors and offsets to enhance network accuracy and simplify fractional bit selection through uniform quantization. On the other hand, non-uniform quantization is well-suited for DNNs with parameters following a normal distribution, helping reduce data width requirements. This paper introduces a novel non-uniform representation called MLB (Multi-Level Binary), which encompasses and extends integer representation. We propose an architecture that objectively compares these representations, demonstrating that MLB is a superset of integer representation in terms of accuracy. Our comprehensive analysis spans various data widths, parallel factorization, and DNN models. Our findings indicate that MLB outperforms integer representation in energy efficiency for lower bit widths (2–5 bits), whereas integer representation is more advantageous for higher bit widths (4–8 bits). Specifically, our work shows an average energy improvement of 1.1 to 2.3× and area saving up to 1.7× compared to integer multiply-accumulate (MAC) units, while preserving network accuracy. This research provides insights into the optimal choice of number representation based on bit-width requirements, highlighting the potential of MLB in enhancing the performance and efficiency of DNNs on embedded devices. Ali Ansarmohammadi, Reza Hojabr, Marzie Mastalizade, Najmeh Nazari, Mostafa E. Salehi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Assessing and Mitigating Heterogeneity-Driven Security Threats in the CloudabstractCloud computing has become crucial for the commercial world due to its computational capacity, storage capabilities, scalability, software integration, and billing convenience. Initially, clouds were relatively homogeneous, but now diverse machine configurations in heterogeneous clouds are recognized for their improved application performance and energy efficiency. This shift is driven by the integration of various hardware to accommodate diverse user applications. However, alongside these advancements, security threats like micro-architectural attacks are increasing concerns for cloud providers and users. Studies like Repttack and Cloak & Co-locate highlight the vulnerability of heterogeneous clouds to co-location attacks, where attacker and victim instances are placed together. The ease of these attacks isn’t solely linked to heterogeneity but also correlates with how heterogeneous the target systems are. Despite this, no numerical metrics exist to quantify cloud heterogeneity. This article introduces the Heterogeneity Score (HeteroScore) to evaluate server setups and instances. HeteroScore significantly correlates with co-location attack security. The article also proposes strategies to balance diversity and security. This study pioneers the quantitative analysis connecting cloud heterogeneity and infrastructure security. Chongzhou Fang, Najmeh Nazari, Behnam Omidi, Han Wang 0020, Aditya Puri, Manish Arora, Setareh Rafatirad, Houman Homayoun, Khaled N. Khasawneh |
ACM Trans. Internet Techn. | 2 |
| 2024 | Architectural Whispers: Robust Machine Learning Models Fingerprinting via Frequency Throttling Side-ChannelsabstractMachine Learning (ML) security practices include hiding ML model architectures to protect intellectual property and prevent attacks. We introduce a novel fingerprinting attack using frequency throttling-based Side-Channel Attack (SCA) to detect an ML model's architecture family by converting power side-channel data into timing variations. This method involves using adversary kernels and a time series ML classifier to discern the architecture from execution time patterns during model operation. We achieved up to 96% accuracy in identifying known ML models' architecture families under Ring 0 privileges and we demonstrated its effectiveness across different platforms. Moreover, our code is publicly available 1. Najmeh Nazari, Chongzhou Fang, Hosein Mohammadi Makrani, Behnam Omidi, Mahdi Eslamimehr, Setareh Rafatirad, Avesta Sasan, Hossein Sayadi, Khaled N. Khasawneh, Houman Homayoun |
DAC | 1 |
| 2024 | SpecScope: Automating Discovery of Exploitable Spectre Gadgets on Black-Box MicroarchitecturesabstractTransient execution attacks pose information leakage risks in current systems. Disabling speculative execution, though mitigating the issue, results in significant performance loss. Accurate identification of vulnerable gadgets is essential for balancing security and performance. However, uncovering all covert channels is challenging due to complex microarchitectural analysis. This paper introduces SpecScope, a framework for automating the detection of Spectre gadgets in code using a black-box microarchitecture approach. SpecScope focuses on contention between transient and non-transient instructions to precisely identify and reduce false-positive Spectre gadgets, minimizing mitigation overhead. Tested on public libraries, SpecScope outperforms existing methods, reducing False-Positive rates by 8.9% and increasing True-Positive rates by 10.4%. Najmeh Nazari, Behnam Omidi, Chongzhou Fang, Hosein Mohammadi Makrani, Setareh Rafatirad, Avesta Sasan, Houman Homayoun, Khaled N. Khasawneh |
DATE | 1 |
| 2024 | Securing On-Chip Learning: Navigating Vulnerabilities and Potential Safeguards in Spiking Neural Network ArchitecturesabstractOn-chip learning is the process of training or updating machine learning models directly on specialized hardware. This approach differs from traditional machine learning, which typically conducts training on external computing resources like Central Processing Units (CPUs) or Graphics Processing Units (GPUs). On-chip learning offers several advantages, including reduced latency, improved energy efficiency, enhanced privacy, and adaptability. Consequently, it holds great promise for enabling intelligent decision-making and adaptability in resource-constrained edge and IoT devices while addressing privacy concerns. In Spiking Neural Network (SNN), on-chip learning is enabled by adjusting synaptic weights, allowing the network’s behavior to dynamically align with desired outcomes. However, this adaptability may introduce potential security vulnerabilities. Unmitigated security risks in on-chip learning can lead to various threats, including data leaks, unauthorized access, and even adversarial manipulation of the learning process. This manuscript aims to provide a comprehensive overview of the security risks associated with on-chip learning, with a focus on potential vulnerabilities within the SNN architecture. We will explore real-world scenarios where these vulnerabilities can be exploited and outline protective measures and mitigation strategies to address these security concerns. Najmeh Nazari, Kevin Immanuel Gubbi, Banafsheh S. Latibari, Md Muhtasim Alam Chowdhury, Chongzhou Fang, Avesta Sasan, Setareh Rafatirad, Houman Homayoun, Soheil Salehi |
ISCAS | 1 |
| 2024 | Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
Chongzhou Fang, Ning Miao, Shaurya Srivastav, Jialin Liu 0006, Ruoyu Zhang 0002, Ruijie Fang, Asmita 0001, Ryan Tsang, Najmeh Nazari, Han Wang 0020, Houman Homayoun |
USENIX Security Symposium | 9 |
| 2024 | Forget and Rewire: Enhancing the Resilience of Transformer-based Models against Bit-Flip Attacks
Najmeh Nazari, Hosein Mohammadi Makrani, Chongzhou Fang, Hossein Sayadi, Setareh Rafatirad, Khaled N. Khasawneh, Houman Homayoun |
USENIX Security Symposium | 1 |
| 2023 | Don't Cross Me! Cross-layer System SecurityabstractThe computing landscape has undergone significant transformations in recent decades. Modern computation systems involve multiple layers across software and hardware architecture, exposing various security vulnerabilities that can be exploited by attackers. In this paper, we review security threats in these systems and provide insights into future directions in the topic of cross-layer security. Najmeh Nazari, Chongzhou Fang, Sai Manoj Pudukotai Dinakarrao, Houman Homayoun |
DAC | 1 |
| 2023 | Inter-Layer Hybrid Quantization Scheme for Hardware Friendly Implementation of Embedded Deep Neural NetworksabstractCompression techniques have been widely deployed to amortize the model size and inference computations of Deep Neural Networks (DNNs), particularly for embedded systems. In this work, we propose an inter-layer approach that deploys a weight distribution aware quantization scheme (a hybrid of fixed-point and power-of-two) and multi-precision (3-bit and 4-bit) to better use heterogeneity in FPGA resources. Based on our evaluation, with similar hardware logic and memory resource usage, our proposed approach improved the throughput of the ResNet-18 network by 37% with negligible accuracy loss compared to the state-of-the-art on an embedded FPGA. Najmeh Nazari, Mostafa E. Salehi |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | HeteroScore: Evaluating and Mitigating Cloud Security Threats Brought by Heterogeneity
Chongzhou Fang, Najmeh Nazari, Behnam Omidi, Han Wang 0020, Aditya Puri, Manish Arora, Setareh Rafatirad, Houman Homayoun, Khaled N. Khasawneh |
NDSS | 2 |
| 2022 | Repttack: Exploiting Cloud Schedulers to Guide Co-Location Attacks
Chongzhou Fang, Han Wang 0020, Najmeh Nazari, Behnam Omidi, Avesta Sasan, Khaled N. Khasawneh, Setareh Rafatirad, Houman Homayoun |
NDSS | 3 |
| 2021 | HosNa: A DPC++ Benchmark Suite for Heterogeneous ArchitecturesabstractMost data centers equipped their general-purpose processors with hardware accelerators to reduce power consumption and improve utilization. Hardware accelerators offer highly energy-efficient computation for a wide range of applications; however, their programming is not as efficient as processors. To bridge the gap, Intel developed a cloud-based infrastructure called DevCloud that connects Intel® Xeon® Scalable Processors to GPUs and FPGAs to deliver high compute performance for emerging workloads. DevCloud assists developers with their compute-intensive tasks and provides access to precompiled software optimized for Intel® architecture. To reduce programming complexity and minimize the barriers to adopt new innovative hardware technology, Intel also provided a unified, cross-architecture programming model called oneAPI based on the Data-Parallel C++ (DPC++) language. In this paper, we introduce HosNa, the first DPC++ benchmark suite that can be used for the evaluation of the Intel FPGAs and DPC++ productivity. Moreover, we present the characterization of proposed benchmarks and the evaluation of implemented hardware accelerators in terms of speedup and latency. Najmeh Nazari, Hosein Mohammadi Makrani, Hossein Sayadi, Lawrence Landis, Setareh Rafatirad, Houman Homayoun |
ICCD | 1 |
| 2021 | ELC-ECG: Efficient LSTM Cell for ECG Classification Based on Quantized ArchitectureabstractLong Short-Term Memory (LSTM) is one of the most popular and effective Recurrent Neural Network (RNN) models used for sequence learning in applications such as ECG signal classification. Complex LSTMs could hardly be deployed on resource-limited bio-medical wearable devices due to the huge amount of computations and memory requirements. Binary LSTMs are introduced to cope with this problem. However, naive binarization leads to significant accuracy loss in ECG classification. In this paper, we propose an efficient LSTM cell along with a novel hardware architecture for ECG classification. By deploying 5-level binarized inputs and just 1- level binarization for weights, output, and in-memory cell activations, the delay of one LSTM cell operation is reduced 50x with about 0.004% accuracy loss in comparison with full precision design of ECG classification. Seyed Ahmad Mirsalari, Najmeh Nazari, Seyed Ali Ansarmohammadi, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
ISCAS | 2 |
| 2020 | Multi-level Binarized LSTM in EEG Classification for Wearable DevicesabstractLong Short-Term Memory (LSTM) is widely used in various sequential applications. Complex LSTMs could be hardly deployed on wearable and resourced-limited devices due to the huge amount of computations and memory requirements. Binary LSTMs are introduced to cope with this problem, however, they lead to significant accuracy loss in some applications such as EEG classification which is essential to be deployed in wearable devices. In this paper, we propose an efficient multi-level binarized LSTM which has significantly reduced computations whereas ensuring an accuracy pretty close to full precision LSTM. By deploying 5-level binarized weights and inputs, our method reduces area and delay of MAC operation about $31\times and 27\times$ in 65nm technology, respectively with less than 0.01% accuracy loss. In contrast to many compute-intensive deep-learning approaches, the proposed algorithm is lightweight, and therefore, brings performance efficiency with accurate LSTM-based EEG classification to realtime wearable devices. Najmeh Nazari, Seyed Ahmad Mirsalari, Sima Sinaei, Mostafa E. Salehi, Masoud Daneshtalab |
PDP | 1 |
| 2019 | TOT-Net: An Endeavor Toward Optimizing Ternary Neural NetworksabstractHigh computation demands and big memory resources are the major implementation challenges of Convolutional Neural Networks (CNNs) especially for low-power and resource-limited embedded devices. Many binarized neural networks are recently proposed to address these issues. Although they have significantly decreased computation and memory footprint, they have suffered from accuracy loss especially for large datasets. In this paper, we propose TOT-Net, a ternarized neural network with [-1, 0, 1] values for both weights and activation functions that has simultaneously achieved a higher level of accuracy and less computational load. In fact, first, TOT-Net introduces a simple bitwise logic for convolution computations to reduce the cost of multiply operations. To improve the accuracy, selecting proper activation function and learning rate are influential, but also difficult. As the second contribution, we propose a novel piece-wise activation function, and optimized learning rate for different datasets. Our findings first reveal that 0.01 is a preferable learning rate for the studied datasets. Third, by using an evolutionary optimization approach, we found novel piece-wise activation functions customized for TOT-Net. According to the experimental results, TOT-Net achieves 2.15%, 8.77%, and 5.7/5.52% better accuracy compared to XNOR-Net on CIFAR-10, CIFAR-100, and ImageNet top-5/top-1 datasets, respectively. Najmeh Nazari, Mohammad Loni, Mostafa E. Salehi, Masoud Daneshtalab, Mikael Sjödin |
DSD | 1 |