EDBT 2026 Demo / reviewers in the wild / expert
Abdel-Hameed A. Badawy
dblp:86/6737 · also Hameed Badawy
· DBLP profile ↗
41ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-8027-1449ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 2 first-author · 12 since 2021Computer networks · 11 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SentinelEdge: An Attention-Based Defense for Real-Time Mitigation of Adversarial Thermal Manipulations in System-on-ChipsabstractDynamic Thermal Management (DTM) systems are critical to the reliable operation of Multiprocessor System-on-Chips (MPSoCs), yet remain vulnerable to sophisticated thermal manipulation attacks. These attacks, executed through hardware trojans or privilege escalation, can compromise the integrity of thermal sensors, causing performance degradation, accelerated aging, and catastrophic hardware failure by disabling thermal throttling mechanisms. Existing countermeasures rely on reactive detection methods and conventional machine learning models that fail to capture the complex physics governing thermal systems, including thermal coupling across cores and power-frequency interdependencies, making them ineffective against multi-stage attacks that exploit DTM decision-making logic. This work presents a novel transformer-based defense framework that leverages self-attention mechanisms to model rich, system-wide feature interactions for detecting adversarial thermal manipulations in real time. The proposed hybrid architecture integrates an adaptive pre-filtering with dynamic thresholding to achieve an 83x throughput improvement (22,798 samples/second versus 274.53 samples/second for transformer-only baseline) and nearly 50% lower GPU utilization, enabling deployment on resource-constrained embedded platforms. Comprehensive on-device validation on the NVIDIA Jetson AGX Orin board demonstrates substantial thermal regulation improvements, reducing average peak temperatures from 103 °C to 98.5 °C while maintaining a model active ratio of only 2.73%. The framework incorporates an adaptive defense system with load-dependent dynamic thresholding that achieves high F1-scores in detecting the thermal attacks discussed in the literature (0.75 to 0.9). This work bridges the critical gap between simulation-based security research and practical embedded system deployment, establishing a new paradigm for lightweight, attention-based anomaly detection in thermally constrained environments. Mehdi Elahi, Mohamed R. Elshamy, Abdel-Hameed A. Badawy, Ahmad Patooghy |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCsabstractEfficient thermal and power management in modern multiprocessor systems-on-chip (MPSoCs) demands accurate power consumption estimation. One of the state-of-theart approaches, Alternative Blind Power Identification (ABPI), theoretically eliminates the dependence on steady-state temperatures, addressing a major shortcoming of previous approaches. However, ABPI performance has remained unverified in actual hardware implementations. In this study, we conducted the first empirical validation of ABPI on commercial hardware using the NVIDIA Jetson Xavier AGX platform. Our findings reveal that, while ABPI provides computational efficiency and independence from steady-state temperature, it exhibits considerable accuracy deficiencies in real-world scenarios. To overcome these limitations, we introduce a novel approach that integrates Custom Physics-Informed Neural Networks (CPINNs) with the underlying thermal model of ABPI. Our approach utilizes a specialized loss function that constrains the data-driven model according to the system's thermal physics, complemented by NSGA-II multi-objective optimization with efficient convergence strategies to balance estimation accuracy and computational cost. In experimental validation, CPINN-ABPI achieves a reduction of 84.7% CPU and 73.9% GPU in the mean absolute error (MAE) relative to ABPI, with the weighted mean absolute percentage error (WMAPE) improving from 47%-81% to ~ 12%. The method maintains real-time performance with 195.3µs of inference time. Extensive evaluation on unseen benchmarks demonstrates robust generalization without overfitting. Mohamed R. Elshamy, Mehdi Elahi, Ahmad Patooghy, Abdel-Hameed A. Badawy |
IPCCC | 4 |
| 2025 | Gate-Breaker: An LLM-Powered Netlist-to-RTL Reverse Engineering ToolabstractThe escalating sophistication of hardware intellectual property (IP) theft, a multi-billion dollar problem for the semiconductor industry, demands novel approaches to both understanding attack vectors and fortifying defenses. This paper evaluates the potential of Large Language Models (LLMs) to reverse engineer Register Transfer Level (RTL) designs from gate-level netlists. We introduce a framework for netlist-to-RTL conversion, leveraging pattern recognition and code generation capabilities of modern LLMs. Our evaluation of four available LLM models across 156 circuit benchmarks reveals that LLMs can indeed recover functional RTL with reasonable accuracy until the RTL benchmarks become our classified 4th quartile of code complexity. We also provide a similarity metrics-based methodology to evaluate and ascertain the quality of reverse engineering. Among the evaluated models, OpenAI's O3-Mini emerged as the best performer with 75.1 % overall success rate. While performance degrades significantly on the most complex 4th quartile (53.4 % success rate), when O3-Mini does succeed on these challenging designs, it maintains a high Abstract Syntax Tree (AST) similarity of 0.924 and achieves a moderate 0.871 Control Flow Graph (CFG) similarity. Md. Omar Faruque, Peter Jamieson, Ahmad Patooghy, Abdel-Hameed A. Badawy |
IPCCC | 4 |
| 2024 | WIP: Integrating Cybersecurity Education: Implementation of an Undergraduate Course on Malicious Thermal Sensor DefenseabstractContribution: This work-in-progress innovative practice paper presents the inception of an undergraduate course focusing on the Ensemble of Countermeasures for Malicious Thermal Sensors Attacks (ECMTA), representing a novel endeavor at this academic level. Rooted in prior research across academic and industrial domains, this ongoing initiative embodies a journey of exploration and discovery in an emerging field, awaiting feedback from students and our industrial advisory board. Additionally, this proposal advocates for a flipped classroom (FC) model, prioritizing pre-class materials for theoretical understanding and in-class sessions for practical application and collaboration. Simultaneously, this project is an ongoing investigation into the effectiveness of FC methodologies for Hispanic/Latino populations in cybersecurity education. Recognizing the current dearth of consensus in the literature regarding this demographic's response to FC, the project aims to address this gap through a comprehensive assessment of Hispanic/Latino students' attitudes and academic outcomes, with findings expected by the end of 2024. Background: The surge in IoT devices revolutionized industries, offering convenience and connectivity. Yet, they pose substantial cybersecurity challenges, notably in thermal sensor vulnerabilities. Attackers' exploitation of these sensors to manipulate the temperatures of IOT devices underscores the urgent necessity for robust cybersecurity in IoT ecosystems. Intended outcome: This course's intended outcome is multifaceted. Students will grasp thermal sensor vulnerabilities and learn varied countermeasures to mitigate risks. Practical skills will be honed through hands-on lab exercises and simulations. Additionally, they'll foster a mindset of continual learning, which is vital for evolving cybersecurity careers. Following this study, we aim to assess the effectiveness of flipped classroom methodologies for Hispanic/Latino populations in cybersecurity education. Application Design: The course offers theoretical lectures, practical workshops, research projects, and assessments focused on defending against thermal sensor attacks. It integrates cybersecurity into undergraduate curricula, fostering innovative teaching methods. Ultimately, it aims to prepare a new generation of cybersecurity professionals for security challenges. Amin Malek Mohammadi, Ahmad Patooghy, Abdel-Hameed A. Badawy |
FIE | 3 |
| 2024 | HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning ApproachabstractThe growing necessity for enhanced processing capabilities in edge devices with limited resources has led us to develop effective methods for improving high-performance computing (HPC) applications. In this paper, we introduce LASP (Lightweight Autotuning of Scientific Application Parameters), a novel strategy designed to address the parameter search space challenge in edge devices. Our strategy employs a multi-armed bandit (MAB) technique focused on online exploration and exploitation. Notably, LASP takes a dynamic approach, adapting seamlessly to changing environments. We tested LASP with four HPC applications: Lulesh, Kripke, Clomp, and Hypre. Its lightweight nature makes it particularly well-suited for resource-constrained edge devices. By employing the MAB framework to efficiently navigate the search space, we achieved significant performance improvements while adhering to the stringent computational limits of edge devices. Our experimental results demonstrate the effectiveness of LASP in optimizing parameter search on edge devices. Abrar Hossain, Abdel-Hameed A. Badawy, Mohammad A. Islam 0001, Tapasya Patki, Kishwar Ahmed |
HiPC | 2 |
| 2024 | Cluster-BPI: Efficient Fine-Grain Blind Power Identification for Defending against Hardware Thermal Trojans in Multicore SoCsabstractModern multicore System-on-Chips (SoCs) include hardware monitoring mechanisms to measure total power consumption, but these aggregate measurements are insufficient for fine-grained thermal and power management. This paper introduces an improved Clustering Blind Power Identification (ICBPI), an approach to improve the sensitivity and robustness of the Blind Power Identification (BPI) approach, which identifies the power consumption of different cores and the thermal model of an SoC using only thermal sensor measurements and the total power consumption. The proposed approach enhances BPI’s initialization step (specifically the non-negative matrix factorization, which is crucial for BPI accuracy) by incorporating density-based spatial clustering of of noise applications (DBSCAN). This is done to maximize the physical relationship between the temperature and power consumption, ensuring more accurate power estimates. Our simulations demonstrate two tasks to validate the proposed approach. The first evaluates the power accuracy per core on four different multicores, including a heterogeneous processor, showing that ICBPI significantly improves accuracy without overheads. For example, in a four-core SoC, error rates are reduced by 77.56% compared to vanilla BPI and by 68.44% compared to the state-of-the-art approach called BPISS. The second task focuses on enhancing the precision and robustness of the detection and localization of malicious thermal sensor attacks in the heterogeneous processor, demonstrating that ICBPI is capable of enhancing security of multicore SoCs. Mohamed R. Elshamy, Mehdi Elahi, Ahmad Patooghy, Abdel-Hameed A. Badawy |
IPCCC | 4 |
| 2024 | The Seeker's Dilemma: Realistic Formulation and Benchmarking for Hardware Trojan DetectionabstractThis work focuses on advancing the security field in the hardware design space by formally defining the problem of Hardware Trojan (HT) detection. The goal is to model HT detection more closely to the real world, i.e., describing the problem as "The Seeker’s Dilemma" (an extension of Hide&Seek on a graph), where a detecting agent is unaware of whether HTs infect circuits or not. Using this problem formulation, we create a benchmark that consists of a mixture of HT-free and HT-infected restructured circuits while preserving their original functionalities. The restructured circuits are randomly infected by HTs, causing a situation where the defender is uncertain if a circuit is infected. Our innovative dataset will help the community better judge the detection quality of different methods by comparing their success rates in circuit classification. We use our benchmark to evaluate three state-of-the-art HT detection tools to show baseline results for this approach. We use Principal Component Analysis to assess the strength of our benchmark, where we observe that some restructured HT-infected circuits are mapped closely to HT-free circuits, leading to significant label misclassification by detectors. Amin Sarihi, Ahmad Patooghy, Abdel-Hameed A. Badawy, Peter Jamieson |
IPCCC | 3 |
| 2024 | Trojan playground: a reinforcement learning framework for hardware Trojan insertion and detection
Amin Sarihi, Ahmad Patooghy, Peter Jamieson, Abdel-Hameed A. Badawy |
J. Supercomput. | 4 |
| 2024 | Scalable Experimental Bounds for Entangled Quantum State FidelitiesabstractEstimating the state preparation fidelity of highly entangled states on noisy intermediate-scale quantum (NISQ) devices is important for benchmarking and application considerations. Unfortunately, exact fidelity measurements quickly become prohibitively expensive, as they scale exponentially as O (3 N for N -qubit states, using full state tomography with measurements in all Pauli bases combinations. However, Somma et al.established that the complexity could be drastically reduced when looking at fidelity lower bounds for states that exhibit symmetries, such as Dicke states and GHZ states. These bounds must still be tight enough for larger states to provide reasonable estimations on NISQ devices. For the first time and more than 15 years after the theoretical introduction, we report meaningful lower bounds for the state preparation fidelity of all Dicke states up to N =10 and all GHZ states up to N =20 on Quantinuum H1 ion-trap systems using efficient implementations of recently proposed scalable circuits for these states. Our achieved lower bounds match or exceed previously reported exact fidelities on superconducting systems for much smaller states. Furthermore, we provide evidence that for large Dicke states \(\left|\smash{D_{N/2}^{N}} \right\rangle\) , we may resort to a GHZ-based approximate state preparation to achieve better fidelity. This work provides a path forward to benchmarking entanglement as NISQ devices improve in size and quality. Shamminuj Aktar, Andreas Bärtschi, Abdel-Hameed A. Badawy, Stephan J. Eidenbenz |
ACM Trans. Quantum Comput. | 3 |
| 2023 | Scalable Experimental Bounds for Dicke and GHZ States FidelitiesabstractEstimating the state preparation fidelity of highly entangled states on noisy intermediate-scale quantum (NISQ) devices is an important task for benchmarking and application considerations. Unfortunately, exact fidelity measurements quickly become prohibitively expensive, as they scale exponentially as O(3N) for N-qubit states, using full state tomography with measurements in all Pauli bases combinations. However, Somma et al. [20] established that the complexity could be drastically reduced when looking at fidelity lower bounds for states that exhibit symmetries, such as Dicke States and GHZ States. For larger states, these bounds still need to be tight enough to provide reasonable estimations on NISQ devices. Shamminuj Aktar, Abdel-Hameed A. Badawy, Andreas Bärtschi, Stephan J. Eidenbenz |
CF | 2 |
| 2023 | Efficient Intra-Rack Resource Disaggregation for HPC Using Co-Packaged DWDM PhotonicsabstractThe diversity of workload requirements and increasing hardware heterogeneity in emerging high performance computing (HPC) systems motivate resource disaggregation. Resource disaggregation allows compute and memory resources to be allocated individually as required to each workload. However, it is unclear how to efficiently realize this capability and cost-effectively meet the stringent bandwidth and latency requirements of HPC applications. To that end, we describe how modern photonics can be co-designed with modern HPC racks to implement flexible intra-rack resource disaggregation and fully meet the bit error rate (BER) and high escape bandwidth of all chip types in modern HPC racks. Our photonic-based disaggregated rack provides an average application speedup of 11% (46% maximum) for 25 CPU and 61% for 24 GPU benchmarks compared to a similar system that instead uses modern electronic switches for disaggregation. Using observed resource usage from a production system, we estimate that an iso-performance intra-rack disaggregated HPC system using photonics would require 4× fewer memory modules and 2× fewer NICs than a non-disaggregated baseline. George Michelogiannakis, Yehia Arafa, Brandon Cook 0001, Liang Yuan Dai, Abdel-Hameed A. Badawy, Madeleine Glick, Yuyang Wang 0003, Keren Bergman, John Shalf |
CLUSTER | 5 |
| 2023 | BB-ML: Basic Block Performance Prediction using Machine Learning TechniquesabstractRecent years have seen the adoption of Machine Learning (ML) techniques to predict the performance of large-scale applications, mostly at a coarse level. In contrast, we propose to use ML techniques for performance prediction at a much finer granularity, namely at the Basic Block (BB) level, which are single entry, single exit code blocks that are used for analysis by the compilers to break down a large code into manageable pieces. Utilizing ML and BB analysis together can enable scalable hardware-software co-design beyond the current state of the art. In this work, we extrapolate the basic block execution counts of GPU applications and use it for predicting the performance for large input sizes from the counts of smaller input sizes.We trained a Poisson Neural Network (PNN) model using random input values as well as the lowest input values of the application to learn the relationship between inputs and basic block counts. Experimental results show that the model can accurately predict the basic block execution counts of 16 GPU benchmarks. We achieved an accuracy of 93.5% for extrapolating the basic block counts for large input sets when the model is trained using smaller input sets. Additionally, the model shows an accuracy of 97.7% for predicting basic block counts on random instances. In a significant case study, we applied the ML model to CUDA GPU benchmarks for performance prediction across a spectrum of applications, spanning linear algebra to machine learning benchmarks. We employed a diverse set of metrics for evaluation, including global memory requests, tensor cores’ active cycles, and the active cycles of ALU and FMA units. The results from the case study demonstrate that the model is capable of predicting the performance of large datasets with high accuracy. For example, The average error rates for global and shared memory requests are 0.85% and 0.17%, respectively. Furthermore, to address the utilization of the main functional units in Ampere architecture GPUs, we calculated the active cycles for units like tensor cores, ALU, FMA, and FP64 units. Our predictions for the active cycles show an average error of 2.3% for the ALU and 10.66% for the FMA units, while the maximum observed error across all tested applications and units reaches 18.5%. Hamdy Abdelkhalik, Shamminuj Aktar, Yehia Arafa, Atanu Barai, Gopinath Chennupati, Nandakishore Santhi, Nishant Panda, Nirmal Prajapati, Nazmul Haque Turja, Stephan J. Eidenbenz, Abdel-Hameed A. Badawy |
ICPADS | 11 |
| 2023 | Securing Network-on-chips Against Fault-injection and Crypto-analysis Attacks via Stochastic Anonymous RoutingabstractNetwork-on-chip (NoC) is widely used as an efficient communication architecture in multi-core and many-core System-on-chips (SoCs). However, the shared communication resources in an NoC platform, e.g., channels, buffers, and routers, might be used to conduct attacks compromising the security of NoC-based SoCs. Most of the proposed encryption-based protection methods in the literature require leaving some parts of the packet unencrypted to allow the routers to process/forward packets accordingly. This reveals the source/destination information of the packet to malicious routers, which can be exploited in various attacks. For the first time, we propose the idea of secure, anonymous routing with minimal hardware overhead to encrypt the entire packet while exchanging secure information over the network. We have designed and implemented a new NoC architecture that works with encrypted addresses. The proposed method can manage malicious and benign failures at NoC channels and buffers by bypassing failed components with a situation-driven stochastic path diversification approach. Hardware evaluations show that the proposed security solution combats the security threats at the affordable cost of 1.5% area and 20% power overheads chip-wide. Ahmad Patooghy, Mahdi Hasanzadeh, Amin Sarihi, Mostafa Abdelrehim, Abdel-Hameed A. Badawy |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2022 | Hardware Trojan Insertion Using Reinforcement LearningabstractThis paper utilizes Reinforcement Learning (RL) as a means to automate the Hardware Trojan (HT) insertion process to eliminate the inherent human biases that limit the development of robust HT detection methods. An RL agent explores the design space and finds circuit locations that are best for keeping inserted HTs hidden. To achieve this, a digital circuit is converted to an environment in which an RL agent inserts HTs such that the cumulative reward is maximized. Our toolset can insert combinational HTs into the ISCAS-85 benchmark suite with variations in HT size and triggering conditions. Experimental results show that the toolset achieves high input coverage rates (100% in two benchmark circuits) that confirms its effectiveness. Also, the inserted HTs have shown a minimal footprint and rare activation probability. Amin Sarihi, Ahmad Patooghy, Peter Jamieson, Abdel-Hameed A. Badawy |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | Performance Evaluation of an Out-of-Order RISC-V CPU: A SPEC INT 2017 StudyabstractAfter almost a decade of waiting, SPEC CPU 2017 was released in 2017. CPU designers have adopted the new benchmarks to evaluate the performance of their designs. Compared to its predecessor SPEC CPU 2006, the average number of code lines, instructions, and memory operations has significantly increased. In this paper, we contrast SPEC CPU 2017 and SPEC CPU 2006 benchmarks regarding performance metrics on a RISC-V processor. Principal Component Analysis (PCA) results show that, with only a few exceptions, the analyzed SPEC CPU 2017 workloads are a subset of SPEC CPU INT 2006 workloads in terms of RISC-V performance metrics. Although the benchmark subsets are very similar, we identified some outliers that would be interesting to use for microarchitecture studies as complementary workloads to SPEC CPU 2006. Amin Sarihi, Michael A. Schoenfelder, Abdel-Hameed A. Badawy |
IPCCC | 3 |
| 2022 | Low-overhead Hardware Supervision for Securing an IoT Bluetooth-enabled Device: Monitoring Radio Frequency and Supply VoltageabstractOver the past decade, the number of Internet of Things (IoT) devices increased tremendously. In particular, the Internet of Medical Things (IoMT) and the Industrial Internet of Things (IIoT) expanded dramatically. Resource restrictions on IoT devices and the insufficiency of software security solutions raise the need for smart Hardware-Assisted Security (HAS) solutions. These solutions target one or more of the three C’s of IoT devices: Communication, Control, and Computation. Communication is an essential technology in the development of IoT. Bluetooth is a widely-used wireless communication protocol in small portable devices due to its low energy consumption and high transfer rates. In this work, we propose a supervisory framework to monitor and verify the operation of a Bluetooth system-on-chip (SoC) in real-time. To verify the operation of the Bluetooth SoC, we classify its transmission state in real-time to ensure a secure connection. Our overall classification accuracy is measured as 98.7%. We study both power supply current (IVDD) and RF domains to maximize the classification performance and minimize the overhead of our proposed supervisory system. Abdelrahman Elkanishy, Paul M. Furth, Derrick T. Rivera, Abdel-Hameed A. Badawy |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | PPT-Multicore: performance prediction of OpenMP applications using reuse profiles and analytical modeling
Atanu Barai, Yehia Arafa, Abdel-Hameed A. Badawy, Gopinath Chennupati, Nandakishore Santhi, Stephan J. Eidenbenz |
J. Supercomput. | 3 |
| 2021 | Securing network-on-chips via novel anonymous routingabstractNetwork-on-Chip (NoC) is widely used as an efficient communication architecture in multi-core and many-core System-on-Chips (SoCs). However, the shared communication resources in NoCs, e.g., channels, buffers, and routers might be used to conduct attacks compromising the security of NoC-based SoCs. Almost all of the proposed encryption-based protection methods in the literature need to leave some parts of the packet unencrypted to allow the routers to process/forward packets accordingly. This uncovers the source/destination information of the packet to malicious routers, which can be used in various attacks. In this paper, we propose the idea of secure anonymous routing with minimal hardware overhead to hide the source/destination information while exchanging secure information over the network. The proposed method uses a novel source-routing algorithm that works with encrypted destination addresses and prevents malicious routers from discovering the source/destination of secure packets. To support our proposal, we have designed and implemented a new NoC architecture that works with encrypted addresses. The conducted hardware evaluations show that the proposed security solution combats the security threats at an affordable cost of 1% area and 10% power overheads chip-wide. Amin Sarihi, Ahmad Patooghy, Mahdi Hasanzadeh, Mostafa Abdelrehim, Abdel-Hameed A. Badawy |
NOCS | 5 |
| 2021 | Load-Aware Dynamic Time Synchronization in Parallel Discrete Event SimulationabstractTraditional Parallel Discrete Event Simulation (PDES) systems employ a monolithic approach for choosing their thread synchronization protocol. They either implement a Time Window-based conservative synchronization or an optimistic event processing capability based on the Time Warp synchronization. In this paper, we show that this binary choice is suboptimal and unnecessary, particularly in the realistic situation where the load distribution across the simulation domain changes over time. We thus propose a new PDES synchronization scheme, called Hybrid PDES, that dynamically switches between conservative and optimistic synchronization protocols based on the simulation run time characteristics. Ali Eker, Yehia Arafa, Abdel-Hameed A. Badawy, Nandakishore Santhi, Stephan J. Eidenbenz, Dmitry V. Ponomarev |
SIGSIM-PADS | 3 |
| 2021 | Hybrid, scalable, trace-driven performance modeling of GPGPUsabstractIn this paper, we present PPT-GPU, a scalable performance prediction toolkit for GPUs. PPT-GPU achieves scalability through a hybrid high-level modeling approach where some computations are extrapolated and multiple parts of the model are parallelized. The tool primary prediction models use pre-collected memory and instructions traces of the workloads to accurately capture the dynamic behavior of the kernels. Yehia Arafa, Abdel-Hameed A. Badawy, Ammar ElWazir, Atanu Barai, Ali Eker, Gopinath Chennupati, Nandakishore Santhi, Stephan J. Eidenbenz |
SC | 2 |
| 2021 | Joint security and performance improvement in multilevel shared cachesabstractAbstract Multilevel cache architectures are widely used in modern heterogeneous systems for performance improvement. However, satisfying the performance and security requirements at the same time is a challenge for such systems. A simple and efficient timing attack on the shared portions of multilevel hierarchical caches and its corresponding countermeasure is proposed here. The proposed attack prolongs the execution time of the victim threads by inducing intentional race conditions in shared memory spaces. Then, a thread‐mapping algorithm to detect such race conditions between a group of threads and resolve them as a countermeasure against the attack is proposed. The proposed countermeasure dynamically monitors races on cache blocks and distributes existing and new threads on processing cores to minimize cache contention. Upon detection of a high contention rate that might be either due to an attack or a natural race condition, two mechanisms, namely cache access‐rate reduction and thread migration, will be used by the countermeasure algorithm to resolve the race situation. Evaluations on SPECCPU 2006 benchmark suite show that the proposed algorithm not only protects the system against the introduced attack but also boosts the overall system performance by an average of 46.35% and 55.92% for the worst and average cases, respectively. Amin Sarihi, Ahmad Patooghy, Mahdi Amininasab, Mohammad Shokrolah Shirazi, Abdel-Hameed A. Badawy |
IET Inf. Secur. | 5 |
| 2020 | Verified instruction-level energy consumption measurement for NVIDIA GPUsabstractGPUs are prevalent in modern computing systems at all scales. They consume a significant fraction of the energy in these systems. However, vendors do not publish the actual cost of the power/energy overhead of their internal microarchitecture. In this paper, we accurately measure the energy consumption of various PTX instructions found in modern NVIDIA GPUs. We provide an exhaustive comparison of more than 40 instructions for four high-end NVIDIA GPUs from four different generations (Maxwell, Pascal, Volta, and Turing). Furthermore, we show the effect of the CUDA compiler optimizations on the energy consumption of each instruction. We use three different software techniques to read the GPU on-chip power sensors, which use NVIDIA's NVML API and provide an in-depth comparison between these techniques. Additionally, we verified the software measurement techniques against a custom-designed hardware power measurement. The results show that Volta GPUs have the best energy efficiency of all the other generations for the different categories of the instructions. This work should aid in understanding NVIDIA GPUs' microarchitecture. It should also make energy measurements of any GPU kernel both efficient and accurate. Yehia Arafa, Ammar ElWazir, Abdelrahman Elkanishy, Youssef Aly, Ayatelrahman Elsayed, Abdel-Hameed A. Badawy, Gopinath Chennupati, Stephan J. Eidenbenz, Nandakishore Santhi |
CF | 6 |
| 2020 | Fast, accurate, and scalable memory modeling of GPGPUs using reuse profilesabstractIn this paper, we introduce an accurate and scalable memory modeling framework for General Purpose Graphics Processor units (GPGPUs), PPT-GPU-Mem. That is Performance Prediction Tool-Kit for GPUs Cache Memories. PPT-GPU-Mem predicts the performance of different GPUs' cache memory hierarchy (L1 & L2) based on reuse profiles. We extract a memory trace for each GPU kernel once in its lifetime using the recently released binary instrumentation tool, NVBIT. The memory trace extraction is architecture-independent and can be done on any available NVIDIA GPU. PPT-GPU-Mem can then model any NVIDIA GPU caches given their parameters and the extracted memory trace. We model Volta Tesla V100 and Turing TITAN RTX and validate our framework using different kernels from Polybench and Rodinia benchmark suites in addition to two deep learning applications from Tango DNN benchmark suite. We provide two models, MBRDP (Multiple Block Reuse Distance Profile) and OBRDP (One Block Reuse Distance Profile), with varying assumptions, accuracy, and speed. Our accuracy ranges from 92% to 99% for the different cache levels compared to real hardware while maintaining the scalability in producing the results. Finally, we illustrate that PPT-GPU-Mem can be used for design space exploration and for predicting the cache performance of future GPUs. Yehia Arafa, Abdel-Hameed A. Badawy, Gopinath Chennupati, Atanu Barai, Nandakishore Santhi, Stephan J. Eidenbenz |
ICS | 2 |
| 2020 | NVIDIA GPGPUs Instructions Energy ConsumptionabstractIn this work, we accurately measure the energy consumption of the different instructions that can be executed in modern NVIDIA GPGPUs. We use three different software techniques to read the GPU on-chip power sensors, which use NVIDIA's NVML API and provide an in-depth comparison between these techniques. Additionally, we verified the software measurement techniques against a custom-designed hardware power measurement. The results show that Volta GPUs have the best energy efficiency of all the other generations for the different categories of the instructions. This work should give GPU architects and developers a more concrete understanding of these representative NVIDIA GPUs' microarchitecture. It should also make energy measurements of any GPU kernel both efficient and accurate. Yehia Arafa, Ammar ElWazir, Abdelrahman Elkanishy, Youssef Aly, Ayatelrahman Elsayed, Abdel-Hameed A. Badawy, Gopinath Chennupati, Stephan J. Eidenbenz, Nandakishore Santhi |
ISPASS | 6 |
| 2019 | POSTER: GPUs Pipeline Latency AnalysisabstractIn this work, we propose a very low overhead and portable analysis for exposing the hidden latency of each individual instruction executing in the pipeline and different access latencies of the various memory hierarchies at the microarchitecture level. We also show the impact of the possible optimizations a CUDA compiler have over the various latencies. We run our evaluation on seven different high-end NVIDIA GPUs from five different generations/architectures namely: Kepler, Maxwell, Pascal, Volta, and Turing. Yehia Arafa, Abdel-Hameed A. Badawy, Gopinath Chennupati, Nandakishore Santhi, Stephan J. Eidenbenz |
ASAP | 2 |
| 2019 | GPUs Cache Performance Estimation using Reuse Distance AnalysisabstractGPU architects have introduced on-chip memories in GPUs to provide local storage nearby processing to reduce the traffic to the device global memory. From then on-wards, modeling to predict the cache performance has been an active area of research. However, due to the complexities found in this highly parallel hardware, this has not been a straightforward task. In this paper, we propose a memory model to predict the entire cache performance (L1 & L2 caches) in GPUs. Our model is based on reuse distance. We use an analytical probabilistic measure of the reuse distance distributions from the memory traces of an application to predict the hit rates. The application’s memory trace is extracted using NVIDIA’s SASSI instrumentation tool. We use 20 different kernels from Polybench and Rodinia benchmark suites and compare our model to the real hardware. The results show that the average prediction accuracy of the model over all the kernels is 86.7% compared to the real device with higher accuracy for the L2 (95.26%) cache than the L1. Furthermore, extracting the application’s memory trace is on average 4. 9x slower compared to the kernels running without instrumentation. This overhead is much smaller than other published results. Furthermore, our model is very flexible where it takes into account the different cache parameters thus it can be used for design space exploration and sensitivity analysis. Yehia Arafa, Gopinath Chennupati, Atanu Barai, Abdel-Hameed A. Badawy, Nandakishore Santhi, Stephan J. Eidenbenz |
IPCCC | 4 |
| 2018 | Initial Explorations of Sparse Matrix-Vector Multiplication on EMU's Migratory Memory Side ProcessingabstractSparse matrices are common place in modern computing and due to the rise of machine learning algorithms it has become apparent that sparse matrix operations play a large role in simulating and modeling systems. Following Amdahl's and Moore's Law, it is clear that simply adding more resources to a CPU will no longer result in improved performance. It is now desirable to make improvements at the architecture level of CPUs to be able to continue to scale and improve performance on key application areas. In this paper, we explore a new architecture being developed by EMU Technology that implements a distributed shared memory and the ability to migrate threads to data. We use some of the well-known sparse matrix methods to compress and solve sparse matrix-vector multiplication (SpMV) and compare the performance of the EMU system and that of a modern x86 architecture. Our results are encouraging but further investigation is required to optimize our current implementation(s) and explore other algorithms for representing sparsity. Matthew Bredin, Abdel-Hameed A. Badawy, Janice O. McMahon, Shannon K. Kuntz |
IEEE BigData | 2 |
| 2018 | 3D-PIM NoCs with Multiple Subnetworks: A Performance and Power EvaluationabstractThe advances in 3D circuit integration have reignited the idea of processing-in-memory (PIM). In this paper, we evaluate 3D mesh-based network on chip (NoC) for 3D-PIM systems with single and multiple network configurations. We study stacked mesh (S-Mesh), which is a mesh-bus hybrid architecture for 3D NoCs that connects vertically stacked 2D meshes through buses. Previous S-Mesh studies have not addressed the problems and modifications needed at the building blocks of the network. We explain in details the internal structure of the S-Mesh, as well as, the problems and possible solutions of connecting 2D meshes using vertical buses. Also, we evaluate the performance of 3D NoCs via two traffic patterns, one of which is a novel traffic pattern that better measures 3D-PIM systems performance. Finally, we use the Rodinia benchmarks to measure the performance under real workloads. We use DSENT to evaluate the power consumption. Our results show ~ 15% performance improvement for the S-Mesh under zero-load packet latency and ~ 11% lower average packet latency for the Rodinia benchmarks. Also, S-Mesh is the low power configuration with the router static power accountable for 90% of the total network power consumption. Abdel-Hameed A. Badawy, Jesus Gardea, Yuho Jin, Jonathan E. Cook 0001 |
IPCCC | 1 |
| 2018 | A performance study of the time-varying cache behavior: a study on APEX, Mantevo, NAS, and PARSEC
Nafiul Siddique, Patricia Grubel, Abdel-Hameed A. Badawy, Jeanine E. Cook |
J. Supercomput. | 3 |
| 2017 | A Probabilistic Monte Carlo Framework for Branch PredictionabstractBranch prediction is crucial in improving the throughput of microprocessors. It reduces branching stalls in the pipeline, which helps to maintain the instruction execution flow. Of these instructions, conditional branches are non-trivial in determining the microprocessor performance and throughput. Modern microprocessors accurately predict the branches using advanced branch prediction techniques. Appropriately estimating the branch mis-predictions benefits to improve the overall performance of an application through effectively saving the CPU cycles. In general, collecting branch prediction statistics using state-of-the-art simulators is time consuming and not scalable. We present a novel Monte Carlo simulation framework that predicts branch mis-prediction rate. Our framework produces results that suggest that the mis-prediction rates on three scientific applications are similar (with an average difference of 0.3%) to that of a Markov model of a 2-bit saturating branch predictor. Bhargava Kalla, Nandakishore Santhi, Abdel-Hameed A. Badawy, Gopinath Chennupati, Stephan J. Eidenbenz |
CLUSTER | 3 |
| 2017 | Analyzing Hybrid Transactional Memory Performance Using Intel SDEabstractDue to the rapidly increasing use of big data, machines are stressed to provide more computing power at higher energy efficiency while maintaining simpler and more scalable computing paradigms. Transactional Memory (TM) is one such technique that can be used for synchronization instead of conventional locks used in critical sections since it has simpler paradigms, is scalable and has better energy efficiency. We used Intel Software Development Emulator (SDE) tool to collect statistics to explore the performance of both Hardware Transactional Memory (HTM) and Hybrid Transactional Memory (HyTM). Our results shows that our designed Adaptive HyTM have higher commit successes (i.e. abort ratio 42x less) than HTM. Therefore, it takes fewer instructions and thus is more energy efficient. Mohammad Qayum, Abdel-Hameed A. Badawy, Jeanine E. Cook |
CLUSTER | 2 |
| 2017 | Optimizing locality in graph computations using reuse distance profilesabstractThis work tries to answer the question of whether or not we should write code differently when the underlying chip microarchitecture is powered by a multicore processor. We use a set of three graph benchmarks each with three different input problems varying in size and connectivity to characterize the importance of how we partition the problem space among cores and how that partitioning can happen at multiple levels of the cache leading to better performance. We explore a design space represented by different parallelization schemes and different graph partitionings. This provides a large and complex space that we characterize using detailed simulation results to see how much gain we can obtain over a baseline legacy parallelization technique with a partition sized to fit in the L1 cache. We show that the legacy parallelization is not the best alternative in most of the cases and other parallelization techniques perform better. We use a PIN computed reuse distance profile to build an execution time prediction model that rank orders the different combinations of parallelization strategies and partitioning sizes. In some cases the prediction is 100% accurate and in some other cases the prediction projects worse performance than the baseline case. We report the difference between the simulated best performing combination and the PIN predicted ones. The M5 performance simulations show gains of up to 20% relative to the baseline. Our prediction scheme can achieve up to 100% of the best performance gains obtained by M5 and up to 48% on average across all of our benchmarks and input sizes. We have shown a new application for reuse distance profiles-i.e., as a tool for helping program developers and compilers to optimize program performance. Abdel-Hameed A. Badawy, Donald Yeung |
IPCCC | 1 |
| 2017 | Probabilistic Monte Carlo simulations for static branch predictionabstractConditional branch instructions have a significant effect on the microprocessor performance and throughput. Accurate branch prediction is crucial in reducing control hazards and improving microprocessor performance. Modern microprocessors accurately predict the branch outcomes using advanced prediction techniques. Estimating branch mis-prediction rates accurately helps to improve the overall performance by saving CPU cycles and power. In general, we run the application programs on cycle accurate hardware simulators such as GEM5 [4], to collect the branch prediction statistics. This method comes out to be time consuming and is also not scalable. We present a novel Monte Carlo simulation framework that produces the branch prediction rate statically, without actually running the application on the hardware. Our framework mimics the execution behavior of the real hardware. It uses one of the three different branch prediction schemes to calculate the branch prediction statistics. It also comments on the branch prediction rates of individual branches. Results suggest that the conditional prediction rates for four scientific applications are similar to that of results from the GEM5 [4] simulator. Bhargava Kalla, Nandakishore Santhi, Abdel-Hameed A. Badawy, Gopinath Chennupati, Stephan J. Eidenbenz |
IPCCC | 3 |
| 2017 | Can Architecture Design Help Eliminate Some Common Vulnerabilities?abstractAs technology improves in size and the number of smart devices increases, security in personal devices undoubtedly becomes an important aspect of today's life. However, the complexity in hardware and software systems expose vulnerabilities in security. Vulnerabilities may exist in many layers of systems and would require a specific inputs or events to trigger it. Discovery of vulnerabilities require significant time and also system specific knowledge, and even then some are difficult to patch.In this paper, we study open source tools for finding potential vulnerabilities and represent the advantages and disadvantages in their use. We present HardVul, a vulnerability checking tool which can be run on any architecture and reports which vulnerabilities were found from our testbed. Strahinja Trecakov, Casey Tran, Abdel-Hameed A. Badawy, Nafiul Siddique, Jaime Acosta, Satyajayant Misra |
MASS | 3 |
| 2017 | Optimizing thin client caches for mobile cloud computing: : Design space exploration using genetic algorithmsabstractSummary The emergence and rapid spread of interest and use of cloud computing as an accessible and expandable, as needed, computing facility on the go, has a very deep affinity to the proliferation of intelligent mobile devices including smartphones and tablets. Together, these technologies have the potential of not leaving anybody behind when it comes to computing applications whether small and personal or large and organizational, and regardless of geographic boundaries and economical conditions. However, many technical challenges still exist that are still delaying the realization of this dream with the responsiveness and quality needed from the user perspective. In this paper, we examine user requirements for access to the cloud through thin clients, handheld and mobile devices. In light of these requirements we characterize some of the needed research developments particularly in the area of device architecture. We present our work in exploring the cache design space for embedded processors using evolutionary techniques for mobile and thin client processors. We present a heuristic, evolutionary approach (genetic algorithm) to exploration that significantly cuts down on the time and resources, obtaining a near optimal design. We demonstrate the real‐world utility of our tool‐chain—“CERE” (pronounced SIRI) short for (CachE Recommendation Engine)—by rapidly and efficiently designing a cache hierarchy, which maximizes the performance of a web browser navigating to a set of popular websites running on a single ARM core. The goal is to improve the users' experience using web browsers. “CERE” made the right choices, and we were able to observe a 17.1%speedup going from the “best” hierarchy relative to the “worst” hierarchy. We will detail potential future directions as well. Abdel-Hameed A. Badawy, Gabriel Yessin, Vikram K. Narayana, David Mayhew, Tarek A. El-Ghazawi |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | LMStr: Local memory store the case for hardware controlled scratchpad memory for general purpose processorsabstractIn this paper, we present a hardware controlled on-chip memory called Local Memory Store (LMStr) that can be used either solely as a scratchpad or as a combination of scratchpad and cache, storing any variable specified by the programmer or extracted by the compiler. LMStr is different than a traditional scratchpad in that it is hardware-controlled and it stores the same type of variables in a block that is allocated based on availability and demand. In this initial work on LMStr, we focus on identifying the potential for LMStr, namely, the advantages of storing temporary and program variables in blocks in LMStr and comparing the performance against a regular cache. To the best of our knowledge, this is the first work where scratchpad has been used in a generalized way where the focus is on storing temporary and programmer specified variables in blocks. We evaluate LMStr on a micro-benchmark and a set of the mini-applications in the mantevo suite. We simulate LMStr in the Structural Simulation Toolkit (SST) simulator. LMStr provides a 10% reduction in average data movement between on-chip and off-chip memory compared to a traditional cache hierarchy. Nafiul Siddique, Abdel-Hameed A. Badawy, Jeanine E. Cook, David Resnick |
IPCCC | 2 |
| 2016 | Exploiting Hierarchical Locality in Deep Parallel ArchitecturesabstractParallel computers are becoming deeply hierarchical. Locality-aware programming models allow programmers to control locality at one level through establishing affinity between data and executing activities. This, however, does not enable locality exploitation at other levels. Therefore, we must conceive an efficient abstraction of hierarchical locality and develop techniques to exploit it. Techniques applied directly by programmers, beyond the first level, burden the programmer and hinder productivity. In this article, we propose the Parallel Hierarchical Locality Abstraction Model for Execution (PHLAME). PHLAME is an execution model to abstract and exploit machine hierarchical properties through locality-aware programming and a runtime that takes into account machine characteristics, as well as a data sharing and communication profile of the underlying application. This article presents and experiments with concepts and techniques that can drive such runtime system in support of PHLAME. Our experiments show that our techniques scale up and achieve performance gains of up to 88%. Ahmad Anbar, Olivier Serres, Engin Kayraklioglu, Abdel-Hameed A. Badawy, Tarek A. El-Ghazawi |
ACM Trans. Archit. Code Optim. | 4 |
| 2014 | Where should the threads go? Leveraging hierarchical data locality to solve the thread affinity dilemmaabstractWe are proposing a novel framework that amelio-rates locality-aware parallel programming models, by defining a hierarchical data locality model extension. We also propose two hierarchical thread partitioning algorithms. These algorithms synthesize hierarchical thread placement layouts that targets minimizing the program's overall communication costs. We demonstrate the effectiveness of our approach using the NAS Parallel Benchmarks implemented in Unified Parallel C (UPC) using a modified Berkeley UPC Compiler and runtime system. We achieved performance gains of up to 88% in performance by applying the placement layouts our algorithms suggest. Ahmad Anbar, Abdel-Hameed A. Badawy, Olivier Serres, Tarek A. El-Ghazawi |
ICPADS | 2 |
| 2013 | Expectations of computing and other STEM students: A comparison for different Class Levels, or (CSE ≠ STEM - CSE) | course levelabstractStudents begin each new course with a set of expectations. These expectations are formed from their experiences in their major, class level, culture, skills, etc. However, faculty and the students are often not on the same page with respect to expectations even though faculty provide students with course syllabi. It is crucial for faculty to understand students' expectations to maximize students' learning, satisfaction, and success. Furthermore, it would promote classroom transparency. There would be no hidden unstated expectations; disappointments during the course can potentially be minimized. We present the results of a survey focused on understanding student expectations. Specifically, we focus on examining the differences in expectations of the students of Computer Science and Engineering (CSE) courses and non-computing STEM courses. We present our analysis and observations of the results using aggregate data for all students at all class levels. We observe various differences and similarities among the STEM fields. Identifying differences is crucial since many non-computing STEM majors are enrolled in computing courses, especially in the lower level courses. We provide a detailed comparison among sophomore and senior level courses in computing, biology and chemistry courses. We also compare sophomore and senior CSE courses. Finally, we discuss the importance of paying attention to all students' needs and expectations. Armed with this knowledge, faculty members can increase transparency in the classroom, student satisfaction, and possibly student retention. Abdel-Hameed A. Badawy, Karl Robert Bruce Schmitt, Sabrina R. Kramer, Katie M. Hrapczynski, Elise A. Larsen, Andrea A. Andrew, Mara R. Dougherty, Matthew W. Miller, Artesha C. Taylor, Breanne Roberston, Alexis Y. Williams, Spencer A. Benson |
FIE | 1 |
| 2010 | Pre-CAD system for normal mammogram detection using local binary pattern featuresabstractBreast cancer is the second leading cause of cancer deaths in women in the U.S. Two main problems appear to affect the decision of detecting and diagnosing breast cancer: the accuracy of the CAD systems used, and the radiologists' performance in reading mammograms. We aim here to improve CAD system's performance by adding a preprocessing step based on the density of the breast to reduce the false negative rate significantly. Mammograms are divided into two distinct categories according to breast density (fatty, and dense). Three LBP-based features are extracted for each of dense and fatty mammograms. A one-class classifier is used for each tissue-type separately to enhance the performance of the overall classification task. The sensitivity for each tissue type was improved significantly when used separately compared to the sensitivity of existing systems that uses all mammograms regardless of tissue type. Mona Y. Elshinawy, Wael W. Abdelmageed, Abdel-Hameed A. Badawy, Mohamed F. Chouikha |
CBMS | 3 |
| 2001 | Evaluating the impact of memory system performance on software prefetching and locality optimizationsabstractSoftware prefetching and locality optimizations are techniques for overcoming the speed gap between processor and memory. In this paper, we evaluate the impact of memory trends on the effectiveness of software prefetching and locality optimizations for three types of applications: regular scientific codes, irregular scientific codes, and pointer-chasing codes. We find for many applications, software prefetching outperforms locality optimizations when there is sufficient memory bandwidth, but locality optimizations outperform software prefetching under bandwidth-limited conditions. The break-even point (for 1 Ghz processors) occurs at roughly 2.5 GBytes/sec on today's memory systems, and will increase on future memory systems. We also study the interactions between software prefetching and locality optimizations when applied in concert. Naively combining the techniques provides robustness to changes in memory bandwidth and latency, but does not yield additional performance gains. We propose and evaluate several algorithms to better integrate software prefetching and locality optimizations, including a modified tiling algorithm, padding for prefetching, and index prefetching. Abdel-Hameed A. Badawy, Aneesh Aggarwal, Donald Yeung, Chau-Wen Tseng |
ICS | 1 |