Mohammadreza Mohammadi

dblp:318/9379 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ReCQ: Residual Compensation for Quantized Convolutional Neural Networks
abstract
Post-training quantization (PTQ) is widely used to reduce the computational and memory requirements of deep neural networks for deployment on resource-constrained platforms. However, aggressive low-bit quantization often introduces significant accuracy degradation, particularly in deep convolutional architectures. To address this challenge, we propose ReCQ, a lightweight residual compensation framework designed to recover quantization-induced errors in convolutional neural networks (CNNs). ReCQ augments quantized models with small depthwise-separable compensation modules that operate in parallel with existing convolutional blocks, learning residual corrections while preserving the original network structure. We evaluate ReCQ across a diverse set of CNN architectures, including VGG, ResNet, MobileNetV2, and ConvNeXt, representing different design paradigms such as plain convolutional networks, residual networks, depthwise-separable architectures, and modern convolutional backbones. Experimental results show that ReCQ consistently improves the accuracy of quantized models compared to standard PTQ and the recent QwT method while introducing only minimal model size overhead. In many configurations, ReCQ achieves higher accuracy with comparable or even smaller parameter overhead, demonstrating its effectiveness and general applicability across modern CNN architectures.
Mohammadreza Mohammadi, Matthew Grenier, Ramtin Zand
ACM Great Lakes Symposium on VLSI1
2025 CrossNAS: A Cross-Layer Neural Architecture Search Framework for PIM Systems
abstract
In this paper, we propose the CrossNAS framework, an automated approach for exploring a vast, multidimensional search space that spans various design abstraction layers-circuits, architecture, and systems-to optimize the deployment of machine learning workloads on analog processing-in-memory (PIM) systems. CrossNAS leverages the single-path one-shot weight-sharing strategy combined with the evolutionary search for the first time in the context of PIM system mapping and optimization. CrossNAS sets a new benchmark for PIM neural architecture search (NAS), outperforming previous methods in both accuracy and energy efficiency while maintaining comparable or shorter search times.
Md Hasibul Amin, Mohammadreza Mohammadi, Jason D. Bakos, Ramtin Zand
ACM Great Lakes Symposium on VLSI2
2025 PixelPrune: Optimizing AIoT Vision Systems via In-Sensor Segmentation and Adaptive Data Transfer
Mohammadreza Mohammadi, Mehrdad Morsali, Sepehr Tabrizchi, Brendan Reidy, Arman Roohi, Shaahin Angizi, Ramtin Zand
ACM Great Lakes Symposium on VLSI1
2025 Privacy Enhancing Federated Learning for Predicting Energy Consumption in Smart Buildings
abstract
Accurate energy consumption forecasting is critical for optimizing energy usage, lowering operational costs, and encouraging sustainability in smart buildings. Machine learning (ML) has developed as an effective method for energy forecasting, using sensor data to anticipate consumption trends and increase efficiency. However, due to regulations such as GDPR and growing privacy concerns, sharing sensitive energy data with third parties is often prohibited, providing issues for traditional centralized ML techniques. Federated Learning (FL) provides a feasible alternative by allowing for decentralized model training across several buildings without explicitly exchanging raw data. This privacy-preserving strategy enables organizations to jointly train reliable models while retaining data sovereignty. Our experimental results demonstrate that by using the CU-BEMS dataset, both FL and centralized forecasting models perform similarly, with an R2score of ≈87%. Furthermore, FL decreases bandwidth use by limiting data transfers, making it a scalable and economical energy management solution for smart buildings. These findings demonstrate FL’s ability to ensure safe, data-driven decision-making for sustainable energy utilization.
Sima Sinaei, Mohammadreza Mohammadi, David Eklund, Henrik Abrahamsson
IJCNN2
2025 A Random Deep Feature Selection Approach to Mitigate Transferable Adversarial Attacks
abstract
Machine learning and deep learning are transformative forces reshaping our networks, industries, services, and ways of life. However, the susceptibility of these intelligent systems to adversarial attacks remains a significant issue. On the one hand, recent studies have demonstrated the potential transferability of adversarial attacks across diverse models. On the other hand, existing defense mechanisms are vulnerable to advanced attacks or are often limited to certain attack types. This study proposes a random deep feature selection approach to mitigate such transferability and improve the robustness of models against adversarial manipulations. Our approach is designed to strengthen deep models against poisoning (e.g., label flipping) and exploratory (e.g., DeepFool, BIM, FGSM, I-FGSM, L-BFGS, C&W, JSMA, and PGD) attacks that are applied in both the training and testing stages, and Transfer Learning-Based Adversarial Attacks. We consider scenarios involving perfect and semi-knowledgeable attackers. The performance of our approach is evaluated through extensive experiments on the renowned UNSW-NB15 dataset, including both real-world and synthetic data, covering a wide range of modern attack behaviors and benign activities. The results indicate that our approach boosts the effectiveness of the target network to over 80% against labelflipping poisoning attacks and over 60% against all major types of exploratory attacks.
Ehsan Nowroozi, Mohammadreza Mohammadi, Ahmad Rahdari, Rahim Taheri, Mauro Conti
IEEE Trans. Netw. Serv. Manag.2
2024 HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROI
abstract
With the rise of tiny IoT devices powered by machine learning (ML), many researchers have directed their focus toward compressing models to fit on tiny edge devices. Recent works have achieved remarkable success in compressing ML models for object detection and image classification on microcontrollers with small memory, e.g., 512kB SRAM. However, there remain many challenges prohibiting the deployment of ML systems that require high-resolution images. Due to fundamental limits in memory capacity for tiny IoT devices, it may be physically impossible to store large images without external hardware. To this end, we propose a high-resolution image scaling system for edge ML, called HiRISE, which is equipped with selective region-of-interest (ROI) capability leveraging analog in-sensor image scaling. Our methodology not only significantly reduces the peak memory requirements, but also achieves up to 17.7× reduction in data transfer and energy consumption.
Brendan Reidy, Sepehr Tabrizchi, Mohammadreza Mohammadi, Shaahin Angizi, Arman Roohi, Ramtin Zand
DAC3
2024 PRIV-DRIVE: Privacy-Ensured Federated Learning using Homomorphic Encryption for Driver Fatigue Detection
abstract
Context: Detecting fatigue in drivers has become increasingly important for safe driving, especially with the use of more smart devices and Internet-connected vehicles. While sharing data between vehicles can enhance fatigue detection systems, privacy concerns pose significant barriers to this sharing process. We propose a Federated Learning (FL) approach for monitoring fatigue-driven behavior to address these challenges. However, there is a concern that the drivers' private information might be leaked in the FL system. In this paper, we introduce PRIV-DRIVE, a novel approach for privacy-enhanced fatigue detection applications. Our method integrates Paillier homo-morphic encryption (PHE) with a top-k parameter selection technique, bolstering privacy and confidentiality in federated fatigue detection systems. This approach reduces communication and computation overhead while ensuring model accuracy. To the best of our knowledge, this is the first paper to implement PHE in FL setups for fatigue detection applications. We ran several experiments and evaluated the PRIV-DRIVE method. The results show substantial efficiency gains with different HE key sizes, reducing computation time by up to 96% and communication traffic by up to 95%. Importantly, these improvements have minimal impact on accuracy, effectively meeting the requirements of fatigue detection applications.
Sima Sinaei, Mohammadreza Mohammadi, Rakesh Shrestha, Mina Alibeigi, David Eklund
DSD2
2024 Edge-Centric Real-Time Segmentation for Autonomous Underwater Cave Exploration
abstract
This paper addresses the challenge of deploying machine learning (ML)-based segmentation models on edge platforms to facilitate real-time scene segmentation for Autonomous Underwater Vehicles (AUVs) in underwater cave exploration and mapping scenarios. We focus on three ML models-U-Net, CaveSeg, and YOLOv8n-deployed on four edge platforms: Raspberry Pi-4, Intel Neural Compute Stick 2 (NCS2), Google Edge TPU, and NVIDIA Jetson Nano. Experimental results reveal that mobile models with modern architectures, such as YOLOv8n, and specialized models for semantic segmentation, like U-Net, offer higher accuracy with lower latency. YOLOv8n emerged as the most accurate model, achieving a 72.5 Intersection Over Union (IoU) score. Meanwhile, the U-Net model deployed on the Coral Dev board delivered the highest speed at 79.24 FPS and the lowest energy consumption at 6.23 mJ. The detailed quantitative analyses and comparative results presented in this paper offer critical insights for deploying cave segmentation systems on underwater robots, ensuring safe and reliable AUV navigation during cave exploration and mapping missions.
Mohammadreza Mohammadi, Adnan Abdullah, Aishneet Juneja, Ioannis M. Rekleitis, Md Jahidul Islam, Ramtin Zand
ICMLA1
2024 Anomaly detection based on LSTM and autoencoders using federated learning in smart electric grid
abstract
In smart electric grid systems, various sensors and Internet of Things (IoT) devices are used to collect electrical data at substations. In a traditional system, a multitude of energy-related data from substations needs to be migrated to central storage, such as Cloud or edge devices, for knowledge extraction that might impose severe data misuse, data manipulation, or privacy leakage. This motivates to propose anomaly detection system to detect threats and Federated Learning to resolve the issues of data silos and privacy of data. In this article, we present a framework to identify anomalies in industrial data that are gathered from the remote terminal devices deployed at the substations in the smart electric grid system. The anomaly detection system is based on Long Short-Term Memory (LSTM) and autoencoders that employs Mean Standard Deviation (MSD) and Median Absolute Deviation (MAD) approaches for detecting anomalies. We deploy Federated Learning (FL) to preserve the privacy of the data generated by the substations. FL enables energy providers to train shared AI models cooperatively without disclosing the data to the server. In order to further enhance the security and privacy properties of the proposed framework, we implemented homomorphic encryption based on the Paillier algorithm for preserving data privacy. The proposed security model performs better with MSD approach using HE-128 bit key providing 97% F1-score and 98% accuracy for K=5 with low computation overhead as compared with HE-256 bit key.
Rakesh Shrestha, Mohammadreza Mohammadi, Sima Sinaei, Alberto Salcines, David Pampliega, Raul Clemente, Ana Lourdes Sanz, Ehsan Nowroozi, Anders Lindgren
J. Parallel Distributed Comput.2
2024 Resisting Deep Learning Models Against Adversarial Attack Transferability via Feature Randomization
abstract
In the past decades, the rise of artificial intelligence has given us the capabilities to solve the most challenging problems in our day-to-day lives, such as cancer prediction and autonomous navigation. However, these applications might not be reliable if not secured against adversarial attacks. In addition, recent works demonstrated that some adversarial examples are transferable across different models. Therefore, it is crucial to avoid such transferability via robust models that resist adversarial manipulations. In this paper, we propose a feature randomization-based approach that resists eight adversarial attacks targeting deep learning models in the testing phase. Our novel approach consists of changing the training strategy in the target network classifier and selecting random feature samples. We consider the attacker with a Limited-Knowledge and Semi-Knowledge conditions to undertake the most prevalent types of adversarial attacks. We evaluate the robustness of our approach using the well-known UNSW-NB15 datasets that include realistic and synthetic attacks. Afterward, we demonstrate that our strategy outperforms the existing state-of-the-art approach, such as the Most Powerful Attack, which consists of fine-tuning the network model against specific adversarial attacks. Further, we demonstrate the practicality of our approach using the VIPPrint dataset through a comprehensive set of experiments. Finally, our experimental results show that our methodology can secure the target network and resists adversarial attack transferability by over 60%.
Ehsan Nowroozi, Mohammadreza Mohammadi, Pargol Golmohammadi, Yassine Mekdad, Mauro Conti, A. Selcuk Uluagac
IEEE Trans. Serv. Comput.2
2023 Balancing Privacy and Accuracy in Federated Learning for Speech Emotion Recognition
abstract
Context: Speech Emotion Recognition (SER) is a valuable technology that identifies human emotions from spoken language, enabling the development of context-aware and personalized intelligent systems.To protect user privacy, Federated Learning (FL) has been introduced, enabling local training of models on user devices.However, FL raises concerns about the potential exposure of sensitive information from local model parameters, which is especially critical in applications like SER that involve personal voice data.Local Differential Privacy (LDP) has prevented privacy leaks in image and video data.However, it encounters notable accuracy degradation when applied to speech data, especially in the presence of high noise levels.In this paper, we propose an approach called LDP-FL with CSS, which combines LDP with a novel client selection strategy (CSS).By leveraging CSS, we aim to improve the representatives of updates and mitigate the adverse effects of noise on SER accuracy while ensuring client privacy through LDP.Furthermore, we conducted model inversion attacks to evaluate the robustness of LDP-FL in preserving privacy.These attacks involved an adversary attempting to reconstruct individuals' voice samples using the output labels provided by the SER model.The evaluation results reveal that LDP-FL with CSS achieved an accuracy of 65-70%, which is 4% lower than the initial SER model accuracy.Furthermore, LDP-FL demonstrated exceptional resilience against model inversion attacks, outperforming the non-LDP method by a factor of 10.Overall, our analysis emphasizes the importance of achieving a balance between privacy and accuracy in accordance with the requirements of the SER application.
Samaneh Mohammadi, Mohammadreza Mohammadi, Sima Sinaei, Ali Balador, Ehsan Nowroozi, Francesco Flammini, Mauro Conti
FedCSIS2
2023 Facial Expression Recognition at the Edge: CPU vs GPU vs VPU vs TPU
abstract
Facial Expression Recognition (FER) plays an important role in human-computer interactions and is used in a wide range of applications. Convolutional Neural Networks (CNN) have shown promise in their ability to classify human facial expressions, however, large CNNs are not well-suited to be implemented on resource-and energy-constrained IoT devices. In this work, we present a hierarchical framework for developing and optimizing hardware-aware CNNs tuned for deployment at the edge. We perform a comprehensive analysis across various edge AI accelerators including NVIDIA Jetson Nano, Intel Neural Compute Stick, and Coral TPU. Using the proposed strategy, we achieved a peak accuracy of 99.49% when testing on the CK+ facial expression recognition dataset. Additionally, we achieved a minimum inference latency of 0.39 milliseconds and a minimum power consumption of 0.52 Watts.
Mohammadreza Mohammadi, Heath Smith, Lareb Khan, Ramtin Zand
ACM Great Lakes Symposium on VLSI1
2023 Caveline Detection at the Edge for Autonomous Underwater Cave Exploration and Mapping
abstract
This paper explores the problem of deploying machine learning (ML)-based object detection and segmentation models on edge platforms to enable realtime caveline detection for Autonomous Underwater Vehicles (AUVs) used for under-water cave exploration and mapping. We specifically investigate three ML models, i.e., U-Net, Vision Transformer (ViT), and YOLOv8, deployed on three edge platforms: Raspberry Pi-4, Intel Neural Compute Stick 2 (NCS2), and NVIDIA Jetson Nano. The experimental results unveil clear tradeoffs between model accuracy, processing speed, and energy consumption. The most accurate model has shown to be U-Net with an 85.53 F1-score and 85.38 Intersection Over Union (IoU) value. Meanwhile, the highest inference speed and lowest energy consumption are achieved by the YOLOv8 model deployed on Jetson Nano operating in the high-power and low-power modes, respectively. The comprehensive quantitative analyses and comparative results provided in the paper highlight important nuances that can guide the deployment of caveline detection systems on underwater robots for ensuring safe and reliable AUV navigation during underwater cave exploration and mapping missions.
Mohammadreza Mohammadi, Sheng-En Huang, Titon Barua, Ioannis M. Rekleitis, Md Jahidul Islam, Ramtin Zand
ICMLA1
2023 Realtime Facial Expression Recognition: Neuromorphic Hardware vs. Edge AI Accelerators
abstract
The paper focuses on real-time facial expression recognition (FER) systems as an important component in various real-world applications such as social robotics. We investigate two hardware options for the deployment of FER machine learning (ML) models at the edge: neuromorphic hardware versus edge AI accelerators. Our study includes exhaustive experiments providing comparative analyses between the Intel Loihi neuromorphic processor and four distinct edge platforms: Raspberry Pi-4, Intel Neural Compute Stick (NSC), Jetson Nano, and Coral TPU. The results obtained show that Loihi can achieve approximately two orders of magnitude reduction in power dissipation and one order of magnitude energy savings compared to Coral TPU which happens to be the least power-intensive and energy-consuming edge AI accelerator. These reductions in power and energy are achieved while the neuromorphic solution maintains a comparable level of accuracy with the edge accelerators, all within the real-time latency requirements.
Heath Smith, James Seekings, Mohammadreza Mohammadi, Ramtin Zand
ICMLA3
2023 Work in Progress: Real-time Transformer Inference on Edge AI Accelerators
abstract
Transformer models have become a dominant architecture in the world of machine learning. From natural language processing to more recent computer vision applications, Transformers have shown remarkable results and established a new state-of-the-art in many domains. However, this increase in performance has come at the cost of ever-increasing model sizes requiring more resources to deploy. Machine learning (ML) models are used in many real-world systems, such as robotics, mobile devices, and internet of things (IoT) devices, that require fast inference with low energy consumption. For batterypowered devices, lower energy consumption directly translates into longer battery life. To address these issues, several edge AI accelerators have been developed. Among these, the Coral Edge TPU has shown promising results for image classification while maintaining very low energy consumption. Many of these devices, including the Coral TPU, were originally designed to accelerate convolutional neural networks, making deployment of Transformers challenging. Here, we propose a methodology to deploy Transformers on Edge TPU. We provide extensive latency, power, and energy comparisons among the leading edge devices and show that our methodology allows for real-time inference of Transformers while maintaining the lowest power and energy consumption of other edge devices on the market.
Brendan Reidy, Mohammadreza Mohammadi, Mohammed E. Elbtity, Heath Smith, Ramtin Zand
RTAS2
2023 An Adversarial Attack Analysis on Malicious Advertisement URL Detection Framework
abstract
Malicious advertisement URLs pose a security risk since they are the source of cyber-attacks, and the need to address this issue is growing in both industry and academia. Several attempts have been made in recent years for malicious URL detection using machine learning (ML). The most widely used techniques extract linguistic features of URL string to extract features like bag-of-words (BoW) before applying ML model. Existing malicious URL detection techniques require effective manual feature engineering that can handle unseen features and generalise to test data. In this study, we extract a novel set of lexical and Web-scrapped features and employ ML techniques for fraudulent advertisement URL detection. The combination set of six different kinds of features precisely overcomes the obfuscation in fraudulent URL classification. Based on distinct statistical properties, we use twelve differently formatted datasets for detection, prediction and classification task. We extend our prediction analysis for mismatched and unlabelled datasets. For this framework, we analyze the performance of four ML techniques: Random Forest, Gradient Boost, XGBoost and AdaBoost in the detection part. With our proposed method, we achieve a false negative rate up to 0.0037 while maintaining high detection accuracy of 99.63%. Moreover, we employ an unsupervised learning technique for data clustering using the${K}$-Means algorithm for the visual analysis. This paper analyses the vulnerability of decision tree-based models using the limited knowledge attack scenario. We considered the exploratory attack during the test phase and implemented Zeroth Order Optimization adversarial attack on the detection models.
Ehsan Nowroozi, Mohammadreza Mohammadi, Mauro Conti
IEEE Trans. Netw. Serv. Manag.3
2023 Employing Deep Ensemble Learning for Improving the Security of Computer Networks Against Adversarial Attacks
abstract
In the past few years, Convolutional Neural Networks (CNN) have demonstrated promising performance in various real-world cybersecurity applications, such as network and multimedia security. However, the underlying fragility of CNN structures poses major security problems, making them inappropriate for use in security-oriented applications, including computer networks. Protecting these architectures from adversarial attacks necessitates using security-wise architectures that are challenging to attack. In this study, we present a novel architecture based on an ensemble classifier that combines the enhanced security of 1-Class classification (known as 1C) with the high performance of conventional 2-Class classification (known as 2C) in the absence of attacks. Our architecture is referred to as the 1.5-Class (cmb-classifier) classifier and is constructed using a final dense classifier, one 2C classifier (i.e., CNNs), and two parallel 1C classifiers (i.e., auto-encoders). In our experiments, we evaluated the robustness of our proposed architecture by considering eight possible adversarial attacks in various scenarios. We performed these attacks on the 2C and cmb-classifier architectures separately. The experimental results of our study showed that the Attack Success Rate (ASR) of the I-FGSM attack against a 2C classifier trained with the N-BaIoT dataset is 0.9900. In contrast, the ASR is 0.0000 for the cmb-classifier.
Ehsan Nowroozi, Mohammadreza Mohammadi, Erkay Savas, Yassine Mekdad, Mauro Conti
IEEE Trans. Netw. Serv. Manag.2
2022 MRAM-based Analog Sigmoid Function for In-memory Computing
abstract
We propose an analog implementation of the transcendental activation function leveraging two spin-orbit torque magnetoresistive random-access memory (SOT-MRAM) devices and a CMOS inverter. The proposed analog neuron circuit consumes 1.8-27x less power, and occupies 2.5-4931x smaller area, compared to the state-of-the-art analog and digital implementations. Moreover, the developed neuron can be readily integrated with memristive crossbars without requiring any intermediate signal conversion units. The architecture-level analyses show that a fully-analog in-memory computing (IMC) circuit that use our SOT-MRAM neuron along with an SOT-MRAM based crossbar can achieve more than 1.1x, 12x, and 13.3x reduction in power, latency, and energy, respectively, compared to a mixed-signal implementation with analog memristive crossbars and digital neurons. Finally, through cross-layer analyses, we provide a guide on how varying the device-level parameters in our neuron can affect the accuracy of multilayer perceptron (MLP) for MNIST classification.
Md Hasibul Amin, Mohammed E. Elbtity, Mohammadreza Mohammadi, Ramtin Zand
ACM Great Lakes Symposium on VLSI3