Hassan Nassar

dblp:297/4157 · DBLP profile ↗
← Back
26ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0003-1566-8997ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 12 first-author · 26 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Efficient Federated Learning with Low-Rank Updates under Homomorphic Encryption
abstract
Federated Learning has been widely adopted for its ability to collaboratively train models without exposing raw data. However, the server-side aggregation process may still leak sensitive information about client data. Homomorphic Encryption enables privacy-preserving aggregation, but it introduces substantial communication overhead for clients and high computational costs for the server. To address these challenges, we propose HEAL-FL, a federated learning framework that is based on low-rank shared basis vectors across clients. Instead of transmitting full encrypted model updates, clients send only encrypted low-rank coefficients, thereby reducing both communication costs and server-side aggregation overhead. Furthermore, HEAL-FL incorporates a communication-efficient basis update scheme that relies exclusively on homomorphic addition at the server. Our evaluation across various homomorphic encryption schemes shows that HEAL-FL reduces client communication and server aggregation costs, leading to improved efficiency of Federated Learning systems. Notably, these savings translate into up to a significant reduction of 38.6% in total training time compared to conventional homomorphic FedAvg with full model parameter transmission, demonstrating the practical benefits of our approach.
Mohamed Aboelenien Ahmed, Mohamed Alsharkawy, Hassan Nassar, Heba Khdr, Jeferson González-Gómez, Jörg Henkel
DATE3
2026 TrustSeed: Lightweight Attestation Protocol for Ensuring LLM Integrity
abstract
Over the last couple of years, large language models have increasingly been integrated into many computing applications. For privacy preservation, they are now deployed on edge devices. However, these deployments are vulnerable to bit flip attacks and backdoor attacks that compromise the integrity of the model. Traditional remote attestation techniques fail to detect such manipulations due to the large model size and the stealthiness of the attacks.In this paper, we present TrustSeed, a lightweight functional attestation protocol that uses a single inference to ensure large language models’ integrity. TrustSeed verifies integrity by applying deterministic, seed-based modifications to model weights within a Trusted Execution Environment and comparing the last intermediate activations and output distribution against a golden reference on the verifier. This approach prevents precomputed or forged responses, ensuring freshness and unpredictability in each attestation round. Our analysis shows that output distribution and last intermediate activations are effective indicators of integrity. We test TrustSeed against bit-flip, data poisoning, and weight poisoning attacks, reliably detecting even single-bit alterations. Extensive evaluations on edge platforms and an HPC system demonstrate minimal overhead and up to 127× faster attestation compared to state-of-the-art full-model hashing.
Mohamed Alsharkawy, Mohamed Aboelenien Ahmed, Hassan Nassar, Jeferson González-Gómez, Heba Khdr, Osama Abboud, Xun Xiao, Jörg Henkel
DATE3
2026 MIQARA: Mixed-Criticality Queue-based Architecture for Reconfigurable Accelerator Platforms
abstract
Coexistence of safety-critical control functions and besteffort computations in mixed-criticality systems poses a challenge in resource allocation and scheduling, as high-criticality jobs must adhere to strict timing guarantees, while lower-criticality jobs should make effective use of available resources without compromising the system’s safety and predictability. This paper introduces MIQARA1, a mixed-criticality queue-based architecture designed for reconfigurable accelerator platforms. MIQARA efficiently combines software-programmable CPUs with reconfigurable hardware, utilizing a dynamic job pipeline, token-based dependency tracking, and out-of-order scheduling to optimize resource utilization. At the same time, MIQARA has been designed to satisfy real-time constraints. MIQARA is evaluated on four FPGA platforms: the Zed Board, DipForty board, ZCU102 board, all of which have ARM CPUs implemented on chip, and Arty A7 with a RISC-V soft-core processor, representing systems that rely on soft CPUs. Results demonstrate substantial performance gains, particularly in terms of execution speed, flexibility, and adaptability to mixed-criticality workloads. The integration of features such as a streaming network further illustrates MIQARA’s scalability to complex data-intensive applications, making it a compelling solution for embedded mixed-criticality systems. MIQARA requires a hardware overhead of 17.8% and achieves a speedup of up to 4×.
Hassan Nassar, Martin Rapp, Lars Bauer, Mostafa Elshimy, Zeynep Demirdag, Jörg Henkel
DATE1
2026 Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAI
abstract
Artificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps.
Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl
DATE11
2025 Through Fabric: A Cross-world Thermal Covert Channel on TEE-enhanced FPGA-MPSoC Systems
abstract
The ever-evolving computing landscape gets more complex in every moment and the need for heterogeneous compute systems becomes more relevant. As the usability of such systems grew, finding methods for securing them became more relevant. Commercial vendors already introduced Trusted Execution Environments (TEEs) for those systems. TEEs serve the need for isolation, where sensitive data are processed in a secure world, and non-trusted applications are executed in the normal world. In this paper, we introduce Through Fabric: a novel attack against TEE-enhanced FPGA-MPSoCs. We show that existing benign hardware accelerators can be manipulated from the secure world to implement a temperature-based covert channel. We successfully run this attack on a commercial FPGA-MPSoC within the OP-TEE environment without additional access rights. We use an open-source implementation of AES for the accelerator and we reach a transmission speed of 2 bits per second with bit error rate of 1.9% and packet error rate of 4.3%. We are the first to show that a TEE can be bypassed on FPGA-MPSoCs via temperature-based covert channel communication.
Hassan Nassar, Jeferson González-Gómez, Varun Manjunath, Lars Bauer, Jörg Henkel
ASP-DAC1
2025 Special Session - Hardware-Software Co-Design for Machine Learning Systems Made Open-Source
abstract
Chip technologies are crucial for the digital transformation of industry and society. Machine Learning (ML) and Artificial Intelligence (AI) are increasingly shaping both daily life and industrial applications, with AI hardware playing a vital role in enabling efficient and scalable ML deployment. However, significant challenges remain in bridging the gap between ML algorithm development and hardware implementation, particularly for edge ML applications where efficiency, power constraints, and adaptability are critical. In such resource-constrained environments, hardware-software co-design becomes essential to achieve the necessary trade-offs between performance, energy efficiency, and system responsiveness. One of the key bottlenecks in ML hardware development is the lack of seamless integration between ML toolchains and electronic design automation (EDA) tools for hardware synthesis and mapping. Current solutions often require extensive manual optimization and costly proprietary software, limiting accessibility and innovation. Open-source tools can play a transformative role in democratizing ML hardware design, fostering collaboration, and addressing the growing shortage of skilled professionals. This paper covers key aspects of hardware-software co-design for ML systems, such as ML algorithms, hardware design, compiler technologies and system security, with a focus on open-source solutions. We highlight the critical need for open-source toolchains that connect ML model development with hardware synthesis and optimization and present solutions for custom hardware, as well as FPGA accelerators.
Mehdi Baradaran Tahoori, Vincent Meyers, Mahboobe Sadeghipourrudsari, Huashuangyang Xu, Jürgen Becker 0001, Tanja Harbaum, Felix Frombach, Julian Höfer, Georgios Sotiropoulos, Jörg Henkel, Zeynep Demirdag, Heba Khdr, Hassan Nassar, Ulf Schlichtmann, Johannes Geier, Philipp van Kempen, Georg Sigl, Stefan Koegler, Matthias Probst, Jürgen Teich, Frank Hannig, Muhammad Sabih, Batuhan Sesli, Norbert Wehn, Lukas Steiner, Wolfgang Kunz, Mohamed Shelkamy Ali
CODES+ISSS13
2025 Late Breaking Results: Decentralized Voting-Based Attestation for IoT Devices
abstract
Remote Attestation (RA) has become a valuable security service for Internet of Things (IoT) devices, as the security of these devices is often not prioritized during the manufacturing process. However, traditional RA schemes suffer from a single point of failure because they rely on a trusted verifier. To address this issue, we propose a voting-based blockchain attestation protocol that provides a reliable solution by eliminating the single point of failure through distributed verification across all nodes. In addition, it offers a traceable and immutable public history of the attestation results, which can be verified by external auditors at any time. Finally, we verify our proposed protocol on three NVIDIA Jetson embedded devices hosting up to 15 attestation nodes.
Mohamed Alsharkawy, Eren Sönmez, Jeferson González-Gómez, Hassan Nassar, Jörg Henkel
DAC4
2025 Late Breaking Results: The Hidden Risks of Activation Duration in PLPUFs
abstract
The security of Internet of Things (IoT) devices is crucial to protect the vast amounts of data exposed due to their widespread adoption. Authentication is one of the key aspects of IoT security, but it becomes increasingly challenging, especially for resource-constrained devices that require lightweight and efficient solutions. Physical Unclonable Functions (PUFs) have emerged as a promising lightweight solution by using the unique physical properties of Integrated Circuits (ICs). Pseudo Liner Feedback Shift Register PUF (PLPUF) is one of the state-of-the-art implementations known for its flexibility in altering the challenge-response space by changing the activation duration. In this work, we demonstrate that selecting an appropriate activation duration for PLPUF is critical, as improper choices can compromise security. By analyzing the linear dependency between the responses of different PLPUF pairs, our results reveal that predictability can reach up to $96 \%$ when an unsuitable activation duration is chosen.
Mohamed Alsharkawy, Jan Zwerschke, Hassan Nassar, Jeferson González-Gómez, Jörg Henkel
DAC3
2025 Hardware/Software Co-Analysis for Worst Case Execution Time Bounds
abstract
Ensuring that safety-critical systems meet timing constraints is crucial to avoid disastrous failures. To verify that timing requirements are met, a worst-case execution time (WCET) bound is computed. However, traditional WCET tools require a predefined timing model for each target processor, which is not available when using custom instruction set extensions. We introduce a novel approach based on hardware-software coanalysis that employs an instrumented hardware description of the target processor, removing the requirement for a separate timing model. We demonstrate this approach by extending the FemtoRV32 Individua RISC-V processor with a custom instruction set extension and show that it accurately models the timing behavior of the resulting system.
Can Joshua Lehmann, Lars Bauer, Hassan Nassar, Heba Khdr, Jörg Henkel
DATE3
2025 REAP-NVM: Resilient Endurance-Aware NVM-Based PUF Against Learning-Based Attacks
abstract
NVM-based PUFs offer secure authentication and cryptographic applications by exploiting NVMs' MLC to generate diverse, ML-attack-resistant responses. Yet, frequent writes degrade these PUFs, lowering reliability and lifespan. This paper presents a model to assess endurance effects on NVM PUFs, guiding the creation of more robust PUFs. Our novel NVM PUF design enhances endurance by evenly distributing writes, thus mitigating cell stress, achieving a 62x improvement over current solutions while preserving security against learning-based attacks.
Hassan Nassar, Ming-Liang Wei, Chia-Lin Yang, Jörg Henkel, Kuan-Hsun Chen
DATE1
2025 FLARE: Fault Attack Leveraging Address Reconfiguration Exploits in Multi-Tenant FPGAs
Jayeeta Chaudhuri, Hassan Nassar, Dennis Gnad, Jörg Henkel, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty
ETS2
2025 Invited Paper: Hardware-Software Co-Design for Highly Optimized, Customized, and Reliable AI Systems
abstract
Over the past decade, AI has been rapidly integrated into our daily life, coming in every shape and size and working across systems from big clouds to IoT. As a result, AI systems are increasingly requiring enhancements in model efficiency, hardware acceleration, and memory systems to satisfy stringent constraints on efficiency, reliability, and security. However, advancing across these fronts is challenging as compute demand outpaces Moore’s-law efficiency, hardening into an AI compute wall and an AI energy wall. Breaking through requires a unified AI co-design loop that co-optimizes algorithms and hardware, including efficient AI-to-hardware mapping, so that ongoing goals (accuracy, sparsity, latency) align with concrete hardware choices (precision modes, interconnects, memory hierarchies) and AI-specific execution and memory-reuse patterns. This paper details the principal co-design challenges, presents complementary strategies, and outlines a practical roadmap toward highly optimized, efficient, reliable, and secure AI systems.
Jörg Henkel, Mehdi Baradaran Tahoori, Heba Khdr, Hassan Nassar, Vincent Meyers, Deming Chen, Selin Yildirim, Yingbing Huang, Nirmal Saxena, Saurabh Hukerikar, Srivi Dhruvanarayan
ICCAD4
2025 DPReF: Decentralized Key Generation Using Physical-Related Functions
abstract
Physical Unclonable Functions (PUFs) serve as a lightweight source to generate cryptographic keys utilizing the inherent physical device properties, making them particularly suitable for resource-constrained environments such as Internet of Things (IoT) devices. Recently, Physical-Related Functions (PReFs) extended PUFs to enable multiple devices to generate similar keys without the need to exchange or store them, improving security. However, state-of-the-art PReF implementations rely on a Trusted Third Party (TTP) to identify relative challenges, introducing a potential vulnerability if the TTP is compromised. In this work, we propose the first decentralized PReF protocol, removing reliance on the TTP and mitigating associated security risks. The proposed protocol allows relative challenges to be identified directly between devices in a decentralized manner. Additionally, we formalize a mathematical model to estimate the minimum number of devices required to build a network, based on the sizes of the PUF and the shared Challenge-Response Pair (CRP).. We demonstrate the generality of our model by verifying it across different types of state-of-the-art PUFs (Arbiter-based Non-Volatile Memory PUF (ANV-PUF) and Pseudo Linear Feedback Shift Register PUF (PLPUF).). We establish a 128 bit cryptographic key using the proposed protocol that matches the state-of-the-art but in a decentralized manner. Moreover, we prove that our protocol can be used to construct hardware-assisted attestation networks using ANV-PUF and PLPUF implementations with a shared secret of 16 bit that allows for both integrity and identity verification.
Mohamed Alsharkawy, Hassan Nassar, Jeferson González-Gómez, Xun Xiao, Osama Abboud, Jörg Henkel
ACM Trans. Embed. Comput. Syst.2
2025 Timekeepers: ML-Driven SDF Analysis for Power-Wasters Detection in FPGAs
abstract
As the integration of FPGAs into cloud computing platforms accelerates, the risk of fault injection attacks - especially through power-wasting designs - becomes increasingly critical. Malicious tenants can upload FPGA designs that, under specific input stimuli, generate excessive power consumption, jeopardizing the integrity of the shared power delivery network (PDN) and enabling denial-of-service or side-channel attacks. Traditional detection techniques relying on netlist and bitstream analysis struggle with generalization and can be evaded through circuit obfuscation and seemingly benign designs. In contrast to these netlist-based approaches, we introduce Timekeepers, a novel detection method that utilizes Standard Delay Format (SDF) timing data combined with machine learning to detect anomalous power behavior in synthesized FPGA designs. Our method trains a decision tree classifier on SDF files generated from both benign and malicious designs, focusing on timing characteristics such as propagation delays and setup/hold violations to identify power wasters at the primitive level. By abstracting away from circuit connectivity and emphasizing timing patterns, our framework is both scalable and robust across different FPGA architectures. The classifier independently evaluates each FPGA component and aggregates the results using a threshold-based voting system to improve detection granularity and reduce false positives. Timekeepers achieves 99.6% accuracy and demonstrates superior performance compared to state-of-the-art solutions. Furthermore, our approach is platform-agnostic and does not require access to netlists or bitstreams, preserving intellectual property confidentiality while enhancing pre-deployment security checks.
Mohamed Fathy, Hassan Nassar, Mohamed Abdelghany, Jörg Henkel
ACM Trans. Embed. Comput. Syst.2
2024 Hacking the Fabric: Targeting Partial Reconfiguration for Fault Injection in FPGA Fabrics
abstract
FPGAs are now ubiquitous in cloud computing infrastructures and reconfigurable system-on-chip, particularly for AI acceleration. Major cloud service providers such as Amazon and Microsoft are increasingly incorporating FPGAs for specialized compute-intensive tasks within their data centers. The availability of FPGAs in cloud data centers has opened up new opportunities for users to improve application performance by implementing customizable hardware accelerators directly on the FPGA fabric. However, the virtualization and sharing of FPGA resources among multiple users open up new security risks and threats. We present a novel fault attack methodology capable of causing persistent fault injections in partial bitstreams during the process of FPGA reconfiguration. This attack leverages powerwasters and is timed to inject faults into bitstreams as they are being loaded onto the FPGA through the reconfiguration manager, without needing to remain active throughout the entire reconfiguration process. Our experiments, conducted on a Pynq FPGA setup, demonstrate the feasibility of this attack on various partial application bitstreams, such as a neural network accelerator unit and a signal processing accelerator unit.
Jayeeta Chaudhuri, Hassan Nassar, Dennis Gnad, Jörg Henkel, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty
ATS2
2024 HBMorphic: FHE Acceleration via HBM-Enabled Recursive Karatsuba Multiplier on FPGA
abstract
Cloud computing offers advantages such as seamless scalability and speedup of computation. Nevertheless, these benefits come with notable tradeoffs, e.g., processing sensitive data without compromising security. Fully Homomorphic Encryption (FHE) solves this by processing of encrypted data. In this work, we develop an FHE hardware accelerator that uses a custom control interface to maximally utilize the bandwidth of HBM, following the memory access patterns of FHE.
Hassan Nassar, Lars Bauer, Jörg Henkel
FCCM1
2024 Covert-Hammer: Coordinating Power-Hammering on Multi-tenant FPGAs via Covert Channels
abstract
With the rise of AI, end of Moore's law, and the digitization of public services, the demand for accelerated computing is growing. To address this demand, major cloud service providers like Amazon Web Services, Microsoft Azure, and Google Cloud Platform have incorporated FPGA instances into their infrastructure with efficient and adaptable resource allocation models. Interest is increasing in multi-tenant FPGAs, which enable multiple users to utilize FPGA resources concurrently, while the FPGA can be split into smaller sections, one per tenant. Nevertheless, it introduces significant security vulnerabilities. For instance, by configuring a malicious circuit in one tenant's section of the FPGA, attacks that cause faults or crash the entire FPGA become feasible, affecting other tenants. By splitting an FPGA into smaller fractions, a single tenant has less potential to cause catastrophic outcomes. However, in this paper, we propose another threat, which is to perform an attack where several malicious tenants coordinate an attack using an unintended covert channel. We practically verify this possibility and introduce such a synchronized and coordinated voltage drop attack from multiple malicious tenants. For synchronization, the malicious tenants use a voltage-based covert channel. Our results show that the communication is robust reaching less than 1% packet error rate and that the attack is successful and avoids state-of-the-art countermeasures.
Hassan Nassar, Philipp Machauer, Dennis Gnad, Lars Bauer, Mehdi Baradaran Tahoori, Jörg Henkel
FPGA1
2024 Co-Designing NVM-based Systems for Machine Learning and In-memory Search Applications
abstract
With the rapid development of the Internet of Things, machine learning applications on edge devices with limited resources face challenges due to large data scales and irregular memory access patterns. Non-volatile memory (NVM) technologies provide promising solutions by offering larger capacity, low leakage power, and data persistence. In this paper, we discuss the potential of NVM technology in enhancing machine learning applications by improving energy efficiency and reducing latency through in-memory computation and different NVM write modes. The insights from this analysis provide valuable guidance to device researchers and system architects working to develop highperformance systems for machine learning and accelerators in large-scale search applications using NVMs.
Jörg Henkel, Lokesh Siddhu, Hassan Nassar, Lars Bauer, Jian-Jia Chen, Christian Hakert, Tristan Taylan Seidl, Kuan-Hsun Chen, Xiaobo Sharon Hu, Mengyuan Li 0001, Chia-Lin Yang, Ming-Liang Wei
ICCAD3
2024 DoS-FPGA: Denial of Service on Cloud FPGAs via Coordinated Power Hammering
abstract
The adoption of FPGA instances by major cloud service providers (CSPs) reflects the growing demand for accelerated and heterogeneous computing across various applications, e.g., AI. To improve the efficiency, utilization and virtualization, multi-tenant FPGAs allow multiple users to utilize FPGA resources concurrently, with each FPGA partition assigned to a separate tenant. However, this introduces significant security vulnerabilities, such as the potential for attacks by configuring a malicious circuit in one tenant's FPGA partition. One notable vulnerability is disrupting the FPGA's power distribution network, leading to faults or even crashing the entire FPGA, affecting other tenants. Usually, such an attack requires a considerable amount of resources. A naive solution would be splitting an FPGA into smaller fractions to reduce the potential for successful Power-Hammering by individual tenants and enhance the security. However, our paper demonstrates that even with smaller fractions per tenant, attacks can still occur. We propose the threat of coordinated attacks, where malicious tenants use an unintended covert channel between them. We practically validate this threat in a real cloud computing environment by introducing a synchronized and coordinated power-hammering attack from multiple malicious tenants. These tenants synchronize their actions using a voltage-based covert channel. Our results reveal the success of the attack, surpassing state-of-the-art countermeasures and detection mechanisms with a success rate exceeding 90%, compared to 30% for uncoordinated attacks.
Hassan Nassar, Philipp Machauer, Lars Bauer, Dennis Gnad, Mehdi Baradaran Tahoori, Jörg Henkel
ICCAD1
2024 Meta-Scanner: Detecting Fault Attacks via Scanning FPGA Designs Metadata
abstract
With the rise of the big data, processing in the cloud has become more significant. One method of accelerating applications in the cloud is to use field programmable gate arrays (FPGAs) to provide the needed acceleration for the user-specific applications. Multitenant FPGAs are a solution to increase efficiency. In this case, multiple cloud users upload their accelerator designs to the same FPGA fabric to use them in the cloud. However, multitenant FPGAs are vulnerable to low-level denial-of-service attacks that induce excessive voltage drops using the legitimate configurations. Through such attacks, the availability of the cloud resources to the nonmalicious tenants can be hugely impacted, leading to downtime and thus financial losses to the cloud service provider. In this article, we propose a tool for the offline classification to identify which FPGA designs can be malicious during operation by analysing the metadata of the bitstream generation step. We generate and test 475 FPGA designs that include 38% malicious designs. We identify and extract five relevant features out of the metadata provided from the bitstream generation step. Using ten-fold cross-validation to train a random forest classifier, we achieve an average accuracy of 97.9%. This significantly surpasses the conservative comparison with the state-of-the-art approaches, which stands at 84.0%, as our approach detects stealthy attacks undetectable by the existing methods.
Hassan Nassar, Jonas Krautter, Lars Bauer, Dennis Gnad, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Late Breaking Results: Configurable Ring Oscillators as a Side-Channel Countermeasure
abstract
Side-channel attacks are a threat to computing devices. In this work, we propose a novel countermeasure against power analysis side-channel attacks. This countermeasure uses ring oscillators with runtime-configurable chain lengths to generate noise to hide the effects of the secret intermediate values on the device’s power consumption. We develop our countermeasure to be compatible with a state-of-the-art of side-channel-attack detection mechanism. Therefore, our solution does not incur any extra area overhead as it uses a subset of the circuit needed for detection. We evaluate our countermeasure using the test vector leakage assessment test (TVLA test). When our countermeasure is active no side-channel leakage could be detected.
Hassan Nassar, Simon Pankner, Lars Bauer, Jörg Henkel
DAC1
2023 Memory Carousel: LLVM-Based Bitwise Wear Leveling for Nonvolatile Main Memory
abstract
Emerging non-volatile memory yields, alongside many advantages, technical shortcomings, such as reduced cell lifetime. Although many wear-leveling approaches exist to extend the lifetime of such memories, usually a trade-off for the granularity of wear-leveling has to be made. Due to iterative write schemes (repeatedly sense and write), wear-out of memory in certain systems is directly dependent on the written bit value and thus can be highly imbalanced, requiring dedicated bit-wise wear-leveling. Such a bit-wise wear-leveling so far has only be proposed together with a special hardware support. However, if no dedicated hardware solutions are available, especially for commercial off-the-shelf systems with non-volatile memories, a software solution can be crucial for the system lifetime. In this work, we propose entirely software-based bit-wise wearleveling, where the position of bits within CPU words in main memory is rotated on a regular basis. We leverage the LLVM intermediate representation to adjust load and store operations of the application with a custom compiler pass. Experimental evaluation shows that the lifetime by applying local rotation within the CPU word can be extended by a factor of up to 21×. We also show that our method can incorporate with coarser-grained wear-leveling, e.g. on block granularity and assist achievement of higher lifetime improvements.
Nils Hölscher, Christian Hakert, Hassan Nassar, Kuan-Hsun Chen, Lars Bauer, Jian-Jia Chen, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 ANV-PUF: Machine-Learning-Resilient NVM-Based Arbiter PUF
abstract
Physical Unclonable Functions (PUFs) have been widely considered an attractive security primitive. They use the deviations in the fabrication process to have unique responses from each device. Due to their nature, they serve as a DNA-like identity of the device. But PUFs have also been targeted for attacks. It has been proven that machine learning (ML) can be used to effectively model a PUF design and predict its behavior, leading to leakage of the internal secrets. To combat such attacks, several designs have been proposed to make it harder to model PUFs. One design direction is to use Non-Volatile Memory (NVM) as the building block of the PUF. NVM typically are multi-level cells, i.e, they have several internal states, which makes it harder to model them. However, the current state of the art of NVM-based PUFs is limited to ‘weak PUFs’, i.e., the number of outputs grows only linearly with the number of inputs, which limits the number of possible secret values that can be stored using the PUF. To overcome this limitation, in this work we design the Arbiter Non-Volatile PUF (ANV-PUF) that is exponential in the number of inputs and that is resilient against ML-based modeling. The concept is based on the famous delay-based Arbiter PUF (which is not resilient against modeling attacks) while using NVM as a building block instead of switches. Hence, we replace the switch delays (which are easy to model via ML) with the multi-level property of NVM (which is hard to model via ML). Consequently, our design has the exponential output characteristics of the Arbiter PUF and the resilience against attacks from the NVM-based PUFs. Our results show that the resilience to ML modeling, uniqueness, and uniformity are all in the ideal range of 50%. Thus, in contrast to the state-of-the-art, ANV-PUF is able to be resilient to attacks, while having an exponential number of outputs.
Hassan Nassar, Lars Bauer, Jörg Henkel
ACM Trans. Embed. Comput. Syst.1
2022 CaPUF: Cascaded PUF Structure for Machine Learning Resiliency
abstract
With the rise of the Internet of Things (IoT), resource-constrained and power-constrained devices attract more attention. The need for lightweight solutions as alternatives to resource-intensive applications became more urgent. Moreover, as the number of connected devices grew, authenticating them became more challenging. Traditionally, this would be performed by using hash functions and secure memory to store a key, which both come at a high cost. physical unclonable functions (PUFs) emerged as a suitable lightweight alternative to hash functions to authenticate the devices. Using the inherent minute differences between integrated circuits (ICs), they can generate IC-specific responses for input challenges coming from a so-called verifier. Through the years, machine learning (ML) has been used to attack PUFs by modeling them and accurately predicting their response to a given challenge. This stimulated research on ML-resilient PUFs. This resilience came with the significant area and challenge-to-response delay overheads. In this work, we introduce the novel cascaded PUF (CaPUF) and show that it is resilient against state-of-the-art ML-based attacks, i.e., logistic regression (LR) and support vector machines (SVMs). These attacks could not achieve accuracy better than 52% against our CaPUF, which is only as good as flipping a coin. Additionally, our CaPUF requires 89% less area compared to state-of-the-art ML-resilient PUFs.
Hassan Nassar, Lars Bauer, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 TiVaPRoMi: Time-Varying Probabilistic Row-Hammer Mitigation
abstract
Row-Hammering is a challenge for computing systems that use DRAM. It can cause bit flips in a DRAM row by accessing its neighboring rows. Several mitigation techniques on memory controller level were already suggested. The techniques are in two categories: The first category uses static probabilities, which leads to a performance penalty due to a high number of extra row activations. The second category is based on so-called Tabled Counters, which have large hardware requirements and are mostly infeasible to implement. We introduce a novel Row-Hammer mitigation technique that uses time-varying probabilities combined with a relatively small history table. Our technique reduces the number of extra row activations compared to static probabilistic techniques and it demands less storage than Tabled Counters techniques. Compared to state of the art, our technique offers a good compromise that has 9× - 27× reduced storage requirement than Tabled Counters and 6× - 12× fewer activations than probabilistic techniques.
Hassan Nassar, Lars Bauer, Jörg Henkel
DATE1
2021 LoopBreaker: Disabling Interconnects to Mitigate Voltage-Based Attacks in Multi-Tenant FPGAs
abstract
FPGAs are being offered in the cloud as accelerator resources that can be shared among multiple users (i.e. tenants). Recently, various approaches have shown that fault attacks launched from one tenant region to another are possible, leading to timing faults or crashes of the FPGA. It is, therefore, important that malicious tenants are limited in their ability to cause such security problems. So far, the existing countermeasures against such attacks check the configuration bitstreams before they are reconfigured. Such offline approaches have various practical limitations, e.g. they may force the tenants to unveil their design secrets. In this paper, we present LoopBreaker, a novel runtime solution that can disable the entire activity of a malicious tenant region, in order to rapidly stop a potential attack before it results in a crash (i.e. Denial-of-Service). We implemented and tested multiple attack types and found that realistic attacks demand at least 12–26 µs to be successful. A partial reconfiguration to overwrite the malicious tenant region demands 200 µs in our realworld implementation, which is too slow to prevent the attack from leading to a crash. Instead, our proposed LoopBreaker method only needs 1.5 µs to stop a malicious tenant, which makes it the first online approach that can successfully stop challenging voltage drop-based attacks from causing a crash.
Hassan Nassar, Hanna AlZughbi, Dennis Gnad, Lars Bauer, Mehdi Baradaran Tahoori, Jörg Henkel
ICCAD1