Ahmad Patooghy

dblp:25/5234 · DBLP profile ↗
← Back
46ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0003-2647-2797ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 5 first-author · 10 since 2021Security and privacy · 10 · 3 first-author · 4 since 2021Computer networks · 5 · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing thermal attack resilience in Multi-Processor System-on-Chip through synthetic data generation using LLMs
abstract
The security of Multi-Processor System-on-Chip (MPSoC) architectures has become a critical concern due to their widespread integration into modern computing systems. However, the lack of quality data capturing the behavior of these systems under attack scenarios hinders the development of effective countermeasures. In this study, we propose a framework that facilitates the collection and enrichment of thermal datasets using machine learning for MPSoC security analysis against two distinct thermal attack patterns. To address the limitations associated with imbalanced data, the proposed framework leverages Large Language Models (LLMs) to generate synthetic data, incorporating a feedback-driven optimization mechanism to enhance data quality and alignment with real-world distributions, we developed a prompt-based LLM pipeline for schema-conformant tabular synthesis and evaluate its downstream utility by training standard off-the-shelf ML models on the generated tables. The framework employs a novel prompting method to ensure compatibility across various LLMs, thereby optimizing the synthetic data generation process. Experimental evaluations provide empirical evidence substantiating the framework’s efficiency in generating and enriching synthetic data for training machine learning models. The experimental evaluation reveals that the Light Gradient Boosting Machine algorithm achieved the best performance among all tested models, attaining an accuracy of 93% and an F1-score of 92.80% when trained only on synthetic data and evaluated on real data. The findings highlight its potential as a powerful tool for strengthening MPSoC security by improving the robustness and adaptability of machine learning-based security mechanisms.
Md Rahat Khan, Samiul Islam Niloy, Mahdi Hasanzadeh, Ahmad Patooghy, Kasem Khalil
Knowl. Based Syst.4
2026 SentinelEdge: An Attention-Based Defense for Real-Time Mitigation of Adversarial Thermal Manipulations in System-on-Chips
abstract
Dynamic Thermal Management (DTM) systems are critical to the reliable operation of Multiprocessor System-on-Chips (MPSoCs), yet remain vulnerable to sophisticated thermal manipulation attacks. These attacks, executed through hardware trojans or privilege escalation, can compromise the integrity of thermal sensors, causing performance degradation, accelerated aging, and catastrophic hardware failure by disabling thermal throttling mechanisms. Existing countermeasures rely on reactive detection methods and conventional machine learning models that fail to capture the complex physics governing thermal systems, including thermal coupling across cores and power-frequency interdependencies, making them ineffective against multi-stage attacks that exploit DTM decision-making logic. This work presents a novel transformer-based defense framework that leverages self-attention mechanisms to model rich, system-wide feature interactions for detecting adversarial thermal manipulations in real time. The proposed hybrid architecture integrates an adaptive pre-filtering with dynamic thresholding to achieve an 83x throughput improvement (22,798 samples/second versus 274.53 samples/second for transformer-only baseline) and nearly 50% lower GPU utilization, enabling deployment on resource-constrained embedded platforms. Comprehensive on-device validation on the NVIDIA Jetson AGX Orin board demonstrates substantial thermal regulation improvements, reducing average peak temperatures from 103 °C to 98.5 °C while maintaining a model active ratio of only 2.73%. The framework incorporates an adaptive defense system with load-dependent dynamic thresholding that achieves high F1-scores in detecting the thermal attacks discussed in the literature (0.75 to 0.9). This work bridges the critical gap between simulation-based security research and practical embedded system deployment, establishing a new paradigm for lightweight, attention-based anomaly detection in thermally constrained environments.
Mehdi Elahi, Mohamed R. Elshamy, Abdel-Hameed A. Badawy, Ahmad Patooghy
ACM Trans. Embed. Comput. Syst.4
2026 Comparative Evaluation of GPT Models in FHIR Proficiency
abstract
Ensuring interoperability in healthcare data exchange is vital for advancing patient care, and Fast Healthcare Interoperability Resources (FHIR) has emerged as a cornerstone standard in this effort. As healthcare increasingly integrates AI for managing and interpreting complex data, proficiency in FHIR is essential to ensure seamless and reliable interactions with healthcare systems. This study evaluates the FHIR proficiency of Generative Pre-Trained Transformer (GPT) models, which serves as a critical benchmark for applying AI in healthcare. The performance of GPT-3.5, GPT-4.0, and two custom models was assessed in two FHIR examination scenarios using novel metrics, including Token Processing Cost (TPC), Accuracy-Adjusted Token Processing Cost (ATPC), Comprehensive Performance Index (CPI), and Quality-Adjusted Performance Score (QAPS). GPT-4.0 demonstrated superior accuracy and robustness, while custom models such as the “FHIR Interop Expert” showed strengths in domain-specific tasks through effective prompt engineering. Despite these capabilities, none of the models consistently achieved the \(\geq\) 99% accuracy required for high-stakes healthcare applications. The findings underscore the importance of refining domain-specific training and evaluation methods. The proposed metrics provide a replicable framework for assessing AI readiness, offering a foundation for the responsible and effective integration of AI into healthcare workflows.
Tia Pope, Ahmad Patooghy
ACM Trans. Intell. Syst. Technol.2
2025 A Data-Driven Framework for Performance Assessment of SIEM Solutions
Jason M. Green, Mahmoud Nabil 0001, Abdolhossein Sarrafzadeh, Ahmad Patooghy
CRiSIS4
2025 CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
abstract
Efficient thermal and power management in modern multiprocessor systems-on-chip (MPSoCs) demands accurate power consumption estimation. One of the state-of-theart approaches, Alternative Blind Power Identification (ABPI), theoretically eliminates the dependence on steady-state temperatures, addressing a major shortcoming of previous approaches. However, ABPI performance has remained unverified in actual hardware implementations. In this study, we conducted the first empirical validation of ABPI on commercial hardware using the NVIDIA Jetson Xavier AGX platform. Our findings reveal that, while ABPI provides computational efficiency and independence from steady-state temperature, it exhibits considerable accuracy deficiencies in real-world scenarios. To overcome these limitations, we introduce a novel approach that integrates Custom Physics-Informed Neural Networks (CPINNs) with the underlying thermal model of ABPI. Our approach utilizes a specialized loss function that constrains the data-driven model according to the system's thermal physics, complemented by NSGA-II multi-objective optimization with efficient convergence strategies to balance estimation accuracy and computational cost. In experimental validation, CPINN-ABPI achieves a reduction of 84.7% CPU and 73.9% GPU in the mean absolute error (MAE) relative to ABPI, with the weighted mean absolute percentage error (WMAPE) improving from 47%-81% to ~ 12%. The method maintains real-time performance with 195.3µs of inference time. Extensive evaluation on unseen benchmarks demonstrates robust generalization without overfitting.
Mohamed R. Elshamy, Mehdi Elahi, Ahmad Patooghy, Abdel-Hameed A. Badawy
IPCCC3
2025 Gate-Breaker: An LLM-Powered Netlist-to-RTL Reverse Engineering Tool
abstract
The escalating sophistication of hardware intellectual property (IP) theft, a multi-billion dollar problem for the semiconductor industry, demands novel approaches to both understanding attack vectors and fortifying defenses. This paper evaluates the potential of Large Language Models (LLMs) to reverse engineer Register Transfer Level (RTL) designs from gate-level netlists. We introduce a framework for netlist-to-RTL conversion, leveraging pattern recognition and code generation capabilities of modern LLMs. Our evaluation of four available LLM models across 156 circuit benchmarks reveals that LLMs can indeed recover functional RTL with reasonable accuracy until the RTL benchmarks become our classified 4th quartile of code complexity. We also provide a similarity metrics-based methodology to evaluate and ascertain the quality of reverse engineering. Among the evaluated models, OpenAI's O3-Mini emerged as the best performer with 75.1 % overall success rate. While performance degrades significantly on the most complex 4th quartile (53.4 % success rate), when O3-Mini does succeed on these challenging designs, it maintains a high Abstract Syntax Tree (AST) similarity of 0.924 and achieves a moderate 0.871 Control Flow Graph (CFG) similarity.
Md. Omar Faruque, Peter Jamieson, Ahmad Patooghy, Abdel-Hameed A. Badawy
IPCCC3
2025 Privacy Playground: A Framework for Assessment of Data Anonymization Methods
abstract
Configuring privacy-preserving methods for real-world datasets remains a significant challenge, as optimal parameters vary substantially across different data characteristics and use cases, requiring careful balancing between privacy protection and data utility. Current approaches rely heavily on manual parameter tuning, leading to inefficient and potentially suboptimal privacy-utility trade-offs. To address this critical gap, we present a novel framework that automatically assesses privacypreserving techniques across diverse datasets to set them at the right working spots. Our framework employs statistical analysis to narrow the parameter search space, dramatically reducing the computational overhead typically associated with finding optimal privacy configurations while maintaining high utility. The framework enables data scientists to rapidly identify optimal privacy-utility trade-offs for any privacy-preserving method. We demonstrate the framework's effectiveness by evaluating Differential Privacy (DP) and K-Anonymity (KA) methods on a wide range of input data, i.e., bank, disease, and oil well datasets. Through our experimentation, we observed that DP achieves optimal performance at moderate privacy budgets$(\epsilon$values between 0.01 and 0.0358), while KA's effectiveness varies with dataset characteristics, favoring lower$k$values for smaller datasets and higher$k$values for larger ones.
Bedeabasi John, Mahmoud Mahmoud, Abdolhossein Sarrafzadeh, Ahmad Patooghy
JCC4
2025 Probing AlphaFold's Input Attack Surface via Red-Teaming
abstract
AlphaFold has revolutionized protein structure prediction, enabling near-experimental accuracy directly from amino acid sequences. However, as the model becomes increasingly integrated into synthsetic biology workflows and generative design pipelines, its assumptions around input integrity and trust remain critically underexamined. Safe deployment considerations also require further scrutiny. In this work, we present the first red-teaming study of AlphaFold focused on its input attack surface. We systematically evaluate the model’s behavior across three key input vectors—FASTA sequences, custom structural templates, and multiple sequence alignments (MSAs)—to assess their susceptibility to adversarial misuse, metadata injection, and semantic misalignment. Our findings reveal that AlphaFold lacks robust safeguards against malformed, non-biological, or dual-use-relevant inputs. Harmful sequences are accepted without flagging, auxiliary files can subtly bias predictions, and user-supplied metadata is inconsistently handled. We also identify inconsistencies in AlphaFold’s confidence scoring and reproducibility under standard configurations. We offer a preliminary input risk scorecard characterizing trust failures across interfaces to support future risk assessments. Our findings emphasize the need for lightweight, trust-aware safeguards in structure prediction pipelines and demonstrate how red teaming can help surface critical assumptions in bio-AI systems.
Tia Pope, Ahmad Patooghy
PST2
2025 Privacy-Preserving Machine Learning in IoT: A Study of Data Obfuscation Methods
Yonan Yonan, Mohammad O. Abdullah, Felix Nilsson, Mahdi Fazeli, Ahmad Patooghy, Slawomir Nowaczyk
SECRYPT5
2024 WIP: Integrating Cybersecurity Education: Implementation of an Undergraduate Course on Malicious Thermal Sensor Defense
abstract
Contribution: This work-in-progress innovative practice paper presents the inception of an undergraduate course focusing on the Ensemble of Countermeasures for Malicious Thermal Sensors Attacks (ECMTA), representing a novel endeavor at this academic level. Rooted in prior research across academic and industrial domains, this ongoing initiative embodies a journey of exploration and discovery in an emerging field, awaiting feedback from students and our industrial advisory board. Additionally, this proposal advocates for a flipped classroom (FC) model, prioritizing pre-class materials for theoretical understanding and in-class sessions for practical application and collaboration. Simultaneously, this project is an ongoing investigation into the effectiveness of FC methodologies for Hispanic/Latino populations in cybersecurity education. Recognizing the current dearth of consensus in the literature regarding this demographic's response to FC, the project aims to address this gap through a comprehensive assessment of Hispanic/Latino students' attitudes and academic outcomes, with findings expected by the end of 2024. Background: The surge in IoT devices revolutionized industries, offering convenience and connectivity. Yet, they pose substantial cybersecurity challenges, notably in thermal sensor vulnerabilities. Attackers' exploitation of these sensors to manipulate the temperatures of IOT devices underscores the urgent necessity for robust cybersecurity in IoT ecosystems. Intended outcome: This course's intended outcome is multifaceted. Students will grasp thermal sensor vulnerabilities and learn varied countermeasures to mitigate risks. Practical skills will be honed through hands-on lab exercises and simulations. Additionally, they'll foster a mindset of continual learning, which is vital for evolving cybersecurity careers. Following this study, we aim to assess the effectiveness of flipped classroom methodologies for Hispanic/Latino populations in cybersecurity education. Application Design: The course offers theoretical lectures, practical workshops, research projects, and assessments focused on defending against thermal sensor attacks. It integrates cybersecurity into undergraduate curricula, fostering innovative teaching methods. Ultimately, it aims to prepare a new generation of cybersecurity professionals for security challenges.
Amin Malek Mohammadi, Ahmad Patooghy, Abdel-Hameed A. Badawy
FIE2
2024 Cluster-BPI: Efficient Fine-Grain Blind Power Identification for Defending against Hardware Thermal Trojans in Multicore SoCs
abstract
Modern multicore System-on-Chips (SoCs) include hardware monitoring mechanisms to measure total power consumption, but these aggregate measurements are insufficient for fine-grained thermal and power management. This paper introduces an improved Clustering Blind Power Identification (ICBPI), an approach to improve the sensitivity and robustness of the Blind Power Identification (BPI) approach, which identifies the power consumption of different cores and the thermal model of an SoC using only thermal sensor measurements and the total power consumption. The proposed approach enhances BPI’s initialization step (specifically the non-negative matrix factorization, which is crucial for BPI accuracy) by incorporating density-based spatial clustering of of noise applications (DBSCAN). This is done to maximize the physical relationship between the temperature and power consumption, ensuring more accurate power estimates. Our simulations demonstrate two tasks to validate the proposed approach. The first evaluates the power accuracy per core on four different multicores, including a heterogeneous processor, showing that ICBPI significantly improves accuracy without overheads. For example, in a four-core SoC, error rates are reduced by 77.56% compared to vanilla BPI and by 68.44% compared to the state-of-the-art approach called BPISS. The second task focuses on enhancing the precision and robustness of the detection and localization of malicious thermal sensor attacks in the heterogeneous processor, demonstrating that ICBPI is capable of enhancing security of multicore SoCs.
Mohamed R. Elshamy, Mehdi Elahi, Ahmad Patooghy, Abdel-Hameed A. Badawy
IPCCC3
2024 The Seeker's Dilemma: Realistic Formulation and Benchmarking for Hardware Trojan Detection
abstract
This work focuses on advancing the security field in the hardware design space by formally defining the problem of Hardware Trojan (HT) detection. The goal is to model HT detection more closely to the real world, i.e., describing the problem as "The Seeker’s Dilemma" (an extension of Hide&Seek on a graph), where a detecting agent is unaware of whether HTs infect circuits or not. Using this problem formulation, we create a benchmark that consists of a mixture of HT-free and HT-infected restructured circuits while preserving their original functionalities. The restructured circuits are randomly infected by HTs, causing a situation where the defender is uncertain if a circuit is infected. Our innovative dataset will help the community better judge the detection quality of different methods by comparing their success rates in circuit classification. We use our benchmark to evaluate three state-of-the-art HT detection tools to show baseline results for this approach. We use Principal Component Analysis to assess the strength of our benchmark, where we observe that some restructured HT-infected circuits are mapped closely to HT-free circuits, leading to significant label misclassification by detectors.
Amin Sarihi, Ahmad Patooghy, Abdel-Hameed A. Badawy, Peter Jamieson
IPCCC2
2024 Trojan playground: a reinforcement learning framework for hardware Trojan insertion and detection
Amin Sarihi, Ahmad Patooghy, Peter Jamieson, Abdel-Hameed A. Badawy
J. Supercomput.2
2023 Securing IoT-Based Healthcare Systems Against Malicious and Benign Congestion
abstract
The Internet of Things (IoT) has made it possible to gather patient data through a network of sensors, referred to as the wireless body area network (WBAN). However, the variability of the wireless channels can pose a challenge to the real-time functionality of Medical IoT (MIoT) systems. These delays can occur due to either natural or malicious congestion in the wireless channels. To address this issue, we present an efficient algorithm that partitions the WBAN nodes within an MIoT system to balance the traffic load across the entire system. In our model, each network partition includes an access point (AP) responsible for managing all the WBANs within its coverage range. Our proposed algorithm enables the APs to dynamically readjust their coverage range based on the overall traffic load, thereby evenly distributing the load among APs to alleviate congested areas. Moreover, the APs continuously monitor the traffic to detect and mitigate congested areas, regardless of whether the congestion is due to a natural load or a malicious traffic injection attack. Based on the simulations done using NS2, the proposed algorithm: 1) can resolve congestion of both cases very efficiently; 2) improves the network delay variation by at least 40%; and 3) improves the network energy consumption and network delay by at least 30% and 44%, respectively.
Meisam Kamarei, Ahmad Patooghy, Ahmad Alsharif, Ali Abdullah S. AlQahtani
IEEE Internet Things J.2
2023 Securing Network-on-chips Against Fault-injection and Crypto-analysis Attacks via Stochastic Anonymous Routing
abstract
Network-on-chip (NoC) is widely used as an efficient communication architecture in multi-core and many-core System-on-chips (SoCs). However, the shared communication resources in an NoC platform, e.g., channels, buffers, and routers, might be used to conduct attacks compromising the security of NoC-based SoCs. Most of the proposed encryption-based protection methods in the literature require leaving some parts of the packet unencrypted to allow the routers to process/forward packets accordingly. This reveals the source/destination information of the packet to malicious routers, which can be exploited in various attacks. For the first time, we propose the idea of secure, anonymous routing with minimal hardware overhead to encrypt the entire packet while exchanging secure information over the network. We have designed and implemented a new NoC architecture that works with encrypted addresses. The proposed method can manage malicious and benign failures at NoC channels and buffers by bypassing failed components with a situation-driven stochastic path diversification approach. Hardware evaluations show that the proposed security solution combats the security threats at the affordable cost of 1.5% area and 20% power overheads chip-wide.
Ahmad Patooghy, Mahdi Hasanzadeh, Amin Sarihi, Mostafa Abdelrehim, Abdel-Hameed A. Badawy
ACM J. Emerg. Technol. Comput. Syst.1
2023 ReNo: novel switch architecture for reliability improvement of NoCs
Zahra Shirmohammadi, Yassin Allivand, Fereshte Mozafari, Ahmad Patooghy, Mona Jalal, Sanaz Kazemi Abharian
J. Supercomput.4
2023 Correction to: ReNo: novel switch architecture for reliability improvement of NoCs
Zahra Shirmohammadi, Yassin Allivand, Fereshte Mozafari, Ahmad Patooghy, Mona Jalal, Sanaz Kazemi Abharian
J. Supercomput.4
2022 Hardware Trojan Insertion Using Reinforcement Learning
abstract
This paper utilizes Reinforcement Learning (RL) as a means to automate the Hardware Trojan (HT) insertion process to eliminate the inherent human biases that limit the development of robust HT detection methods. An RL agent explores the design space and finds circuit locations that are best for keeping inserted HTs hidden. To achieve this, a digital circuit is converted to an environment in which an RL agent inserts HTs such that the cumulative reward is maximized. Our toolset can insert combinational HTs into the ISCAS-85 benchmark suite with variations in HT size and triggering conditions. Experimental results show that the toolset achieves high input coverage rates (100% in two benchmark circuits) that confirms its effectiveness. Also, the inserted HTs have shown a minimal footprint and rare activation probability.
Amin Sarihi, Ahmad Patooghy, Peter Jamieson, Abdel-Hameed A. Badawy
ACM Great Lakes Symposium on VLSI2
2022 A multi-application approach for synthesizing custom network-on-chips
Somayeh Kashi, Ahmad Patooghy, Dara Rahmati, Mahdi Fazeli
J. Supercomput.2
2021 Securing network-on-chips via novel anonymous routing
abstract
Network-on-Chip (NoC) is widely used as an efficient communication architecture in multi-core and many-core System-on-Chips (SoCs). However, the shared communication resources in NoCs, e.g., channels, buffers, and routers might be used to conduct attacks compromising the security of NoC-based SoCs. Almost all of the proposed encryption-based protection methods in the literature need to leave some parts of the packet unencrypted to allow the routers to process/forward packets accordingly. This uncovers the source/destination information of the packet to malicious routers, which can be used in various attacks. In this paper, we propose the idea of secure anonymous routing with minimal hardware overhead to hide the source/destination information while exchanging secure information over the network. The proposed method uses a novel source-routing algorithm that works with encrypted destination addresses and prevents malicious routers from discovering the source/destination of secure packets. To support our proposal, we have designed and implemented a new NoC architecture that works with encrypted addresses. The conducted hardware evaluations show that the proposed security solution combats the security threats at an affordable cost of 1% area and 10% power overheads chip-wide.
Amin Sarihi, Ahmad Patooghy, Mahdi Hasanzadeh, Mostafa Abdelrehim, Abdel-Hameed A. Badawy
NOCS2
2021 Joint security and performance improvement in multilevel shared caches
abstract
Abstract Multilevel cache architectures are widely used in modern heterogeneous systems for performance improvement. However, satisfying the performance and security requirements at the same time is a challenge for such systems. A simple and efficient timing attack on the shared portions of multilevel hierarchical caches and its corresponding countermeasure is proposed here. The proposed attack prolongs the execution time of the victim threads by inducing intentional race conditions in shared memory spaces. Then, a thread‐mapping algorithm to detect such race conditions between a group of threads and resolve them as a countermeasure against the attack is proposed. The proposed countermeasure dynamically monitors races on cache blocks and distributes existing and new threads on processing cores to minimize cache contention. Upon detection of a high contention rate that might be either due to an attack or a natural race condition, two mechanisms, namely cache access‐rate reduction and thread migration, will be used by the countermeasure algorithm to resolve the race situation. Evaluations on SPECCPU 2006 benchmark suite show that the proposed algorithm not only protects the system against the introduced attack but also boosts the overall system performance by an average of 46.35% and 55.92% for the worst and average cases, respectively.
Amin Sarihi, Ahmad Patooghy, Mahdi Amininasab, Mohammad Shokrolah Shirazi, Abdel-Hameed A. Badawy
IET Inf. Secur.2
2021 An energy efficient synthesis flow for application specific SoC design
Somayeh Kashi, Ahmad Patooghy, Dara Rahmati, Mahdi Fazeli
Integr.2
2020 Scan-based attack tolerance with minimum testability loss: a gate-level approach
abstract
Scan chain is an architectural solution to facilitate in‐field tests and debugging of digital chips, however, it is also known as a source of security problems, e.g. scan‐based attacks in the chips. The authors conduct a comprehensive gate‐level security analysis on crypto‐chips, which are equipped with a scan chain, and then propose a set of protection mechanisms to immune vulnerable nets of the chips against scan‐based attacks. After extracting the set of most vulnerable nets, they perform net pruning algorithms on them, and gate‐level protection mechanisms to block the information leaking from the nets during test mode. The protection mechanisms employ net masking, net flipping, and net shuffling based on the specifications of every net, i.e. gate‐type driving the net, fan‐out of the net, and net's logical depth. Their evaluations on the hardware‐implemented advanced encryption standard (AES) and data encryption standard (DES) encryption algorithms show 100% for all types of scan‐based attack tolerance, while the area overhead is at most 1.5%, 6.1% for AES and DES crypto‐chip, respectively. As they find the smallest set of nets that have a high contribution to the scan attack, the test coverage loss of their protection mechanism is evaluated to be <0.8%.
Mohammad Taherifard, Mahdi Fazeli, Ahmad Patooghy
IET Inf. Secur.3
2020 Addressing a New Class of Reliability Threats in 3-D Network-on-Chips
abstract
Network-on-chips (NoCs) are vulnerable to transient and permanent faults caused by thermal violations, aging effects, component wear out, or even transient fault sources. Although some of these faults are addressed by previous research, we show that there are reliability threats in 3-D NoCs that go beyond the reliability issues investigated in 2-D interconnect networks. First, we highlight one such class of reliability threats and discuss their manifestations in 3-D NoCs. Second, we propose a thermal, reliability, and performance-aware routing algorithm to tackle: 1) previously established fault models and 2) the new highlighted class of reliability threats in partially connected 3-D NoCs. The proposed routing algorithm takes into account the states of routers and both the horizontal and through silicon via (TSV) links, along with the temperatures of routers and cores. It then routes the packets around failed or overheated links and routers, achieving lower latencies by avoiding misrouting. To achieve this, the proposed routing algorithm uses the concept of vertical link announcement to inform nodes in the network of the working condition of vertical links. We evaluate the proposed routing algorithm under a wide range of working conditions using the access Noxim NoC simulator. Results show that the proposed routing algorithm: 1) is able to tolerate almost any number and pattern of vertical link failures; 2) is reliable against the newly identified reliability threats; and 3) improves the latency and temperature distribution of the network compared to previously proposed routing algorithms.
Ebadollah Taheri, Mihailo Isakov, Ahmad Patooghy, Michel A. Kinsy
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 RaceR: A Thread Mapping Algorithm for Race Reduction in Multi-Level Shared Caches
abstract
Multi-level hierarchical cache architectures are now being widely used in the design and fabrication of multi and many-core chips. However, when two or more threads race to write their own data into the shared-cache, contentions may happen. This natural conflict seriously aggravates the performance of multi-core systems by showing variant performance in multiple runs of even a same program. In this paper, an efficient thread-mapping algorithm is proposed to minimize the cache race condition between threads of multi-core systems. The proposed algorithm, dynamically monitors races on cache blocks and distributes existing and new threads on cores such that the cache contention is minimized. The proposed algorithm uses instructions per cycle (IPC) parameter to detect conflicting threads on the multi-core system. Upon detection of a high contention rate, two mechanisms of cache access rate reduction, and thread migration are used to resolve the race situation. The first solution is a short term one with negligible performance loss, while the former totally resolves the problem with a relatively higher performance cost. Evaluations of the proposed algorithm are done by the use of AKULA simulator alongside SPEC CPU 2006 benchmark suit. Simulation results show that the proposed algorithm improves system performance by average of 6.12% for the SPEC CPU 2006 benchmark suit.
Pezhman Shojaa Sahneh, Amin Sarihi, Benjamin Warburton, Ahmad Patooghy
PDP4
2019 3DEP: A Efficient Routing Algorithm to Evenly Distribute Traffic Over 3D Network-on-Chips
abstract
Due to high manufacturing cost of Through Silicon Via (TSV) in 3D Network-on-Chips (NoCs), not every router is vertically connected. In most 3D NoCs only a subset of TSVs are contrived which results into incomplete 3D NoCs in the vertical dimension. This irregularity introduces new complexity in the design of efficient routing algorithms for partially connected 3D NoCs. In this paper, we propose an efficient routing algorithm to evenly distribute traffic in incomplete 3D NoCs. The proposed algorithm uses turn model analysis to categorize layers, columns, and rows of the NoC into different groups. Then, specific turns are prohibited in each group such that the whole routing is deadlock free, livelock free, and independent of the location of TSVs over the network. By finding the best combination for prohibited turns over the network, we limited the number of required virtual channels to two virtual channel per each physical channel. Simulation results show that the proposed partially adaptive routing has improved packet latency by 32.8% in comparison with Elevator-First algorithm. The advantage of the proposed algorithm is that it shows more improvements on the packet latency and network throughput when the size of network grows.
Fatemeh Vahdatpanah, Mahdi Elahi, Somayeh Kashi, Ebadollah Taheri, Ahmad Patooghy
PDP5
2019 An Efficient Technique to Detect Stealthy Hardware Trojans Independent of the Trigger Size
Seyed Mohammad Sebt, Ahmad Patooghy, Hakem Beitollahi
J. Electron. Test.2
2018 Vulnerability modelling of crypto-chips against scan-based attacks
abstract
In this study, a gate‐level vulnerability model is proposed to detect the potential security holes of crypto‐chips against scan‐based attacks. The proposed model offers a relative measure so‐called vulnerability factor (VF) for each net of a given crypto‐chip. Nets with the highest VFs are considered as the most vulnerable nets of the crypto‐chip. The VF of each gate output is calculated considering (i) VFs of the gate inputs, and (ii) the probability of having a signal transition at the gate output. In order to validate the proposed model, the authors implemented the iterative and pipelined AES, as well as the iterative DES encryption algorithms to find their most vulnerable nets. Then the most vulnerable nets of each design, have been masked by a simple mechanism to explore the accuracy of the proposed model. Results of scan‐based attacks which are done by ModelSim simulations show that by masking only 32, 64 and 32 nets in iterative Advanced Encryption Standard (AES), pipelined AES and iterative Data Encryption Standard(DES) designs, respectively, all of the done attacks are failed. Achieved results of the proposed model in comparison with the signal activity and random approaches demonstrate the superiority of the proposed model.
Mohammad Taherifard, Ahmad Patooghy, Mahdi Fazeli
IET Inf. Secur.2
2017 Crosstalk Free Coding Systems to Protect NoC Channels against Crosstalk Faults
abstract
Reliability of modern multicore and many-core chips is tightly coupled with the reliability of their on-chip networks. Communication channels in current Network-on-Chips (NoCs) are extremely susceptible to crosstalk faults. In this work, we propose a set of rules for generating classes of crosstalk free coding systems to protect communication channels in NoCs against crosstalk faults. Codewords generated through these rules are free of '101' and '010' bit patterns, which are the main sources of crosstalk faults in NoC communication channels. The proposed rules determine: (1) the weights of different bit positions in a coding system to reach crosstalk free codings, and (2) how the coding might be utilized in an NoC to prevent crosstalk generating bit patterns in NoC channels. Using the proposed set of rules, designers can obtain coding systems which are crosstalk free for any widths of communication channels. Compared to conventional Forbidden Pattern Free (FPF) systems, the proposed methodology is able to provide unique representation to any input values at the lower bound of the codeword lengths. Analyses show that the proposed rules, along with the proposed encoding/decoding mechanisms, are effective in preventing forbidden pattern coding systems for network-on-chips of any arbitrary channel width.
Kimia Soleimani, Ahmad Patooghy, Nasim Soltani, Lake Bu, Michel A. Kinsy
ICCD2
2017 3D-AMAP: A Latency-Aware Task Mapping onto 3D Mesh-Based NoCs with Partially-Filled TSVs
abstract
This paper proposes a latency-aware task mapping algorithm called 3D-AMAP for 3D mesh-based NoCs with partially-filled TSVs. The 3D-AMAP algorithm divides communications of a given application graph into Low-volume (LV) and High-Volume (HV) communications. The 3D-AMAP algorithm bypasses the LV communications to partition the given application graph to some subgraphs. Then, 3D-AMAP algorithm fairly assigns 4-neighbor cores of the mesh topology between the high traffic rate tasks of the application graph to reach the bestmapping. The proposed mapping algorithm maps application subgraphs one by one based on their total intra communications considering where the vertical channels are located in the network. Evaluations of the 3D-AMAP mapping algorithm are done in a wide range of working conditions using Access Noxim NoC simulator in terms of network latency. Results show that 3D-AMAP algorithm offers at least 5% and at most 76% in network latencywith respect to NMAP algorithm.
Hesamedin Ziaeeziabari, Ahmad Patooghy
PDP2
2017 Fault-tolerant routing methodology for hypercube and cube-connected cycles interconnection networks
Hossein Habibian, Ahmad Patooghy
J. Supercomput.2
2016 CirKet: A Performance Efficient Hybrid Switching Mechanism for NoC Architectures
abstract
In this paper, we propose and evaluate a hybrid switching mechanism for Network-on-Chips (NoCs). We propose the use of pseudo circuit-switching along with packet-switching in NoC routers. To do this, packets traversing NoC channels are categorized into high and low priority packets which are routed using pseudo circuit and packet switching respectively. Each output port of NoC routers are equipped with a one-bit flag register indicating that the traversing packet is either of low or high priority packet. Using pseudo circuit switching, high priority packets reserve the intermediate routers till the tail flit passes the router. The proposed switching mechanism offers its highest efficiency for applications in which the traffic is dominated by streams. We have used Booksim2 which is a cycle accurate NoC simulator to evaluate the proposed switching technique. Results show that the proposed switching mechanism 1) offers at least 10% improvement in overall flit latency in all tested scenarios as compared with a packet-switched NoC, 2) reduces the power dissipation in the router logic by at least 4%, and 3) imposes a negligible area overhead to NoC router circuitry.
Mohamad FallahRad, Ahmad Patooghy, Hesamedin Ziaeeziabari, Ebadollah Taheri
DSD2
2016 Hardware enlightening: No where to hide your Hardware Trojans!
abstract
IC design and manufacturing chains show steadily growing complexity which provides different third party roles in between. Reprobate parties can take the opportunity to steal a client's IP or insert their malicious circuits-Hardware Trojans-in the original client's design and trigger them in case of need. Trojans are usually inserted in the most hidden internal signals with the lowest activity which increase their chance for not being activated and revealed by clients or end-users. In this paper we propose a method to reduce the number of signals with low activity and hence the chance of inserting hidden trojans. This method is based on an enhanced Logic Encryption approach and uses a 128-bit key. Encryption can also secure the design against IP piracy. Simulation results show that the proposed method can eliminate 83.17% of low activity signals in the circuit.
Seyyed Mohammad Saleh Samimi, Ehsan Aerabi, Zahra Kazemi, Mahdi Fazeli, Ahmad Patooghy
IOLTS5
2016 Phase Change Memory lifetime enhancement via online data swapping
Marzieh Ranjbar Pirbasti, Mahdi Fazeli, Ahmad Patooghy
Integr.3
2011 Numeral-Based Crosstalk Avoidance Coding to Reliable NoC Design
abstract
This paper proposes a Numeral-Based Crosstalk Avoidance Coding (NB-CAC) to protect communication channels of Network-on-Chips (NoCs) against crosstalk faults. The NB-CAC scheme produces code words without bit patterns '101' and '010' to eliminate harmful transition patterns from NoC channels. This is done by the use of a new numeral system proposed in the paper. Using the proposed numeral system, the NB-CAC scheme 1) can be utilized in NoC channels with any arbitrary width, and 2) can be implemented with low area, power, and timing overheads. VHDL and SPICE simulations have been carried out for a wide range of channel widths to evaluate delay, area, and power consumption of the NB-CAC codecs. Results of simulations reveal that the NB-CAC scheme completely removes crosstalk faults from NoC channel. In addition, the NB-CAC scheme provides reductions of 17.3% in area and 31.9% in power-delay product with respect to Fibonacci-based coding which has been recently proposed in literature.
Mansour Shafaei, Ahmad Patooghy, Seyed Ghassem Miremadi
DSD2
2010 An Efficient Method to Reliable Data Transmission in Network-on-Chips
abstract
Data transmission in Network-on-Chips (NoCs) is a serious problem due to cross talk faults happening in adjacent communication links. This paper proposes an efficient flow-control method to enhance the reliability of packet transmission in Network-on-Chips. The method investigates the opposite direction transitions appearing between flits of a packet to reorder the flits in the packet. Flits are reordered in a fixed-size window to reduce: 1) the probability of cross talk occurrence, and 2) the total power consumed for packet delivery. The proposed flow-control method is evaluated by a VHDL-based simulator under different window sizes and various channel widths. Simulation results enable NoC designers to make a trade-off between window size, reliability and power consumption of packet delivery. This method is also compared with other cross talk tolerant methods in terms of reliability and power consumption. Comparison results confirm that the method is a cost efficient solution to overcome the cross talk problem.
Ahmad Patooghy, Hamed Tabkhi, Seyed Ghassem Miremadi
DSD1
2010 Crosstalk modeling to predict channel delay in Network-on-Chips
abstract
Communication channels in Network-on-Chips (NoCs) are highly susceptible to crosstalk faults due to the use of nano-scale VLSI technologies in the fabrication of NoCs. Crosstalk faults cause variable timing delay in NoC channels based on the patterns of transitions appearing on the channels. This paper proposes an analytical model to estimate the timing delay of an NoC channel in the presence of crosstalk faults. The model calculates expected number of 4C, 3C, 2C, and 1C transition patterns to predict delay of a K-bit communication channel. The model is applicable for both non-protected channels and channels which are protected by crosstalk mitigation methods. Spice simulations are done in a wide range of working conditions to validate the proposed model. Delays extracted from the simulations are compared with those obtained from the model. Comparisons show that the proposed model accurately estimates the delay of NoC channels. In addition, the proposed model accelerates the evaluation phase of any crosstalk mitigation method by at least three orders of magnitude.
Ahmad Patooghy, Seyed Ghassem Miremadi, Mansour Shafaei
ICCD1
2010 FiRot: An Efficient Crosstalk Mitigation Method for Network-on-Chips
abstract
This paper proposes an efficient cross talk mitigation method for Network-on-Chips (NoCs). The proposed method investigates flits in each packet to minimize the number of harmful transition patterns appearing on the communication channels of NoC. To do this, the content of every flit is rotated with respect to the previously flit sent through the channel. Rotation is done to find a rotated version of the flit which minimizes the number of harmful transition patterns. A tag field is added into the rotated flit to enable the receiving side to recover the original flit. Maximum number of rotations is bounded by a fixed value to minimize the timing and power overheads of the proposed method. Evaluation of the proposed method is done in both analytical and simulation manners. VHDL-based simulations are carried out for several channel widths and several tag widths. Simulation results confirm that the proposed method effectively overcomes the cross talk problem while its timing and power overheads are negligible. Results of analytical evaluation are also in agreement with the simulation results.
Ahmad Patooghy, Mansour Shafaei, Seyed Ghassem Miremadi, Hajar Falahati, Somayyeh Taheri
PRDC1
2010 A low-overhead and reliable switch architecture for Network-on-Chips
Ahmad Patooghy, Seyed Ghassem Miremadi, Mahdi Fazeli
Integr.1
2009 XYX: A Power & Performance Efficient Fault-Tolerant Routing Algorithm for Network on Chip
abstract
Reliability is one of the main concerns in the design of network on chips due to the use of deep-sub micron technologies in fabrication of such products. This paper proposes a fault-tolerant routing algorithm called XYX which is based on sending redundant packets through the paths with lower traffic loads. The XYX routing algorithm makes a redundant copy of each packet at the source node and exploits two different routing algorithms to route the original and the redundant packets. Since two copies of each packet reach the destination node, the erroneous packet is detected and replaced with the correct one. Due to the use of paths with lower traffic rates for sending redundant packets and minimizing the number of sent redundant packets, the XYX routing algorithm provides lower performance and power overheads as compared to flood-based routing algorithms. Experimental results show that the XYX routing algorithm imposes negligible performance and power consumption overheads while providing almost the same reliability in comparison with flood-based routing algorithms.
Ahmad Patooghy, Seyed Ghassem Miremadi
PDP1
2007 Feedback Redundancy: A Power Efficient SEU-Tolerant Latch Design for Deep Sub-Micron Technologies
abstract
The continuous decrease in CMOS technology feature size increases the susceptibility of such circuits to single event upsets (SEU) caused by the impact of particle strikes on system flip flops. This paper presents a novel SEU-tolerant latch where redundant feedback lines are used to mask the effects of SEUs. The power dissipation, area, reliability, and propagation delay of the presented SEU-tolerant latch are analyzed by SPICE simulations. The results show that this latch consumes about 50% less power and occupies 42% less area than a TMR-latch. However, the reliability and the propagation delay of the proposed latch are still the same as the TMR-latch. the reliability of the proposed latch is also compared with other SEU-tolerant latches.
Mahdi Fazeli, Ahmad Patooghy, Seyed Ghassem Miremadi, Alireza Ejlali
DSN2
2007 Performance Modelling of Necklace Hypercubes
abstract
The necklace hypercube has recently been introduced as an attractive alternative to the well-known hypercube. Previous research on this network topology has mainly focused on topological properties, VLSI and algorithmic aspects of this network. Several analytical models have been proposed in the literature for different interconnection networks, as the most cost-effective tools to evaluate the performance merits of such systems. This paper proposes an analytical performance model to predict message latency in wormhole-switched necklace hypercube interconnection networks with fully adaptive routing. The analysis focuses on a fully adaptive routing algorithm which has been shown to be the most effective for necklace hypercube networks. The results obtained from simulation experiments confirm that the proposed model exhibits a good accuracy under different operating conditions.
Sina Meraji, Hamid Sarbazi-Azad, Ahmad Patooghy
IPDPS3
2007 A Low-Power and SEU-Tolerant Switch Architecture for Network on Chips
abstract
High reliability, high performance, low power consumption are the main objectives in the design of NoCs. These three design objectives are mostly conflicting and should be considered simultaneously in order to have an optimal design. This paper proposes a method based on duplicating the virtual channels of each NoC node as well as parity codes to prevent SEUs from producing erroneous data. The method is compared with two widely used SEU-tolerant methods i.e., the switch to switch and the end to end flow control methods, in terms of reliability, power consumption and performance. A flit level VHDL-based simulator and Synopsys power compiler tool have been used to extract experimental results. The simulation results show the same reliability for all three methods, while the proposed method shows the lowest power consumption and the highest performance almost in all traffic generation rates and all packet error rates.
Ahmad Patooghy, Mahdi Fazeli, Seyed Ghassem Miremadi
PRDC1
2006 Performance Comparison of Partially Adaptive Routing Algorithms
abstract
Partially adaptive routing algorithms are a useful category of routing algorithms due to their simple router logic and restricted adaptivity in selecting the next output channel towards the destination. Several partially adaptive routing algorithms on mesh and hypercube networks have been presented in the literature. But there is no study on evaluating the performance of these algorithms. This paper tries to compare the most important partially adaptive routing algorithms on the mesh and hypercube networks as the most popular topologies for multicomputers. The evaluation has been performed by the use of event driven simulator coded by C++ compiler.
Ahmad Patooghy, Hamid Sarbazi-Azad
AINA (2)1
2006 A Solution to Single Point of Failure Using Voter Replication and Disagreement Detection
abstract
This paper suggests a method, called distributed voting, to overcome the problem of the single point of failure in a TMR system used in robotics and industrial control applications. It uses time redundancy and is based on TMR with disagreement detector feature. This method masks faults occurring in the voter where the TMR system can continue its function properly. The method has been evaluated by injecting faults into Vertex2Pro and Vertex4 Xilinx FPGAs An analytical evolution is also performed. The results of both evaluation approaches show that the proposed method can improve the reliability and the mean time to failure (MTTF) of a TMR system by at least a factor of (2-RV(t)) where RV(t) is the reliability of the voter
Ahmad Patooghy, Seyed Ghassem Miremadi, Abbas Javadtalab, Mahdi Fazeli, Navid Farazmand
DASC1
2006 Analytical performance modelling of partially adaptive routing in wormhole hypercubes
abstract
Although several analytical models have been proposed in the literature for different interconnection networks with different routing algorithms, there is only one work dealing with partially adaptive routing algorithms. This paper proposes an accurate analytical model to predict message latency in wormhole-routed hypercube based networks using the partially adaptive routing algorithm. The results obtained from simulation experiments confirm that the proposed model exhibits a good accuracy for various network sizes and under different operating conditions
Ahmad Patooghy, Hamid Sarbazi-Azad
IPDPS1