VLDB 2026 Research / reviewers in the wild / expert
Nele Mentens
dblp:23/3870
· DBLP profile ↗
76ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0001-8753-7895ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 51 · 5 first-author · 21 since 2021Security and privacy · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Software engineering, systems software and programming languages · 6 · 2 since 2021Computer networks · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | POSTER: Security and Cost Trade-offs of Side-Channel Countermeasures for AES Software on RISC-V SoCsabstractPower side-channel attacks remain a practical concern for cryptographic software deployed on lightweight RISC-V platforms. This poster discusses how several protection strategies influence both leakage resilience and implementation overhead for AES software executed on an FPGA-based RISC-V SoC. Our experiments focus on Tiny-AES running on the Ibex RISC-V core and compare the following protection strategies: software masking, software-generated noise, hardware-generated noise, Secure-Ibex, and CoCo-Ibex. The results show that first-order Correlation Power Analysis succeeds quickly on unprotected SoCs, whereas masking and Secure-Ibex provide the strongest resilience. In contrast, noise-based protection strategies mainly complicate the attack without fully removing exploitable leakage. In addition to security, we compare execution time, code size, hardware usage, power, and energy consumption. Abolfazl Sajadi, Nusa Zidaric, Todor Stefanov, Nele Mentens |
CF | 4 |
| 2026 | Breaking the Scalability Barrier of Content Addressable Memories: A Probabilistic Alternative for Large-Key Associative SearchabstractContent Addressable Memories (CAMs) offer high-speed, deterministic lookups but face significant scalability challenges with large input keys ( \( > \) 100 bits), leading to excessive power, silicon area, and memory costs. This article introduces Probabilistic CAM (P-CAM), a novel architecture designed to overcome these limitations by trading strict determinism for memory efficiency and scalability. P-CAM compresses high-dimensional inputs into fixed-size fingerprints using hashing, making memory requirements independent of key length. P-CAM preserves the constant-time lookup advantage of CAMs, while supporting applications with large keys, such as networking, bioinformatics, and machine learning, where conventional CAMs are impractical. FPGA implementation on Xilinx UltraScale+ devices shows that P-CAM maintains constant query latency and delivers 15 \(\times\) improvement in resource efficiency when handling 384-bit keys, compared to state-of-the-art deterministic CAMs designed for narrower inputs. Although P-CAM’s probabilistic nature introduces a small, controllable false-positive rate, it can be configured for fully deterministic operation under specific constraints. To the best of our knowledge, P-CAM is the first CAM architecture to employ a fingerprint-based probabilistic data structure as the primary storage mechanism for associative lookup, distinguishing it from prior probabilistic approaches that are limited to set membership checks, offering a robust and scalable alternative for modern data-intensive systems. Arish Sateesan, Jo Vliegen, Nele Mentens |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | CarDS - Controller Area Network and Automotive Ethernet Realistic Data SetabstractIntrusion Detection Systems (IDSs) serve as a crucial defense mechanism against cyberattacks targeting the In-Vehicle Network (IVN) of modern, interconnected vehicles. To develop and test new IDS approaches, researchers require realistic IVN data featuring real attacks on moving vehicles. To this end, this paper presents Controller Area Network and Automotive Ethernet Realistic Data Set (CarDS), a novel dataset targeting both the Controller Area Network (CAN) and Automotive Ethernet (AE) traffic of a modern, multi-domain and multi-protocol IVN. Existing datasets are often simulated or limited to basic IVN architectures consisting of only a single CAN bus. Additionally, there are no realistic datasets for AE, despite its growing importance in high-speed in-vehicle communication. CarDS addresses these limitations by providing a labeled, time-synchronized dataset of CAN and AE traces that includes both comprehensive benign profiles and sophisticated attacks. Our traces are captured from an electric vehicle from 2020 featuring a domain-oriented architecture comprising 10 internal CAN buses and 6 AE buses. Specifically, our dataset covers 9h 07m 09s of real IVN data and features 397,383,125 CAN and 180,604,377 AE messages distributed over different scenarios in 258 traces. Wouter Hellemans, Jannis Hamborg, Timm Lauser, Md Masoom Rabbani, Bart Preneel, Christoph Krauß, Nele Mentens |
ACSAC | 7 |
| 2025 | TraceFormer: A Transformer-Based Method for Weight Extraction from AIMC TilesabstractAnalogIn-Memory Computing (AIMC) has emerged as a promising solution to address the performance and energy efficiency limitations of conventional von Neumann architectures for machine learning (ML) applications. This promising approach relies on analog-to-digital converters (ADCs) to enable the integration of AIMC tiles into larger digital systems. In this paper, we investigate the vulnerability of AIMC tiles to power side-channel attacks targeting these ADCs. Specifically, we demonstrate that the numerical values of weights stored in an AIMC tile, that are often a critical asset of an ML model, can be extracted by analyzing the power consumption of the ADCs. With this objective in mind, we propose TraceFormer, which is a two-phase method: 1) we train a Transformer neural network (NN) model to translate captured ADC power traces into digital output values; 2) we utilize a novel input-controlled weight isolation technique in order to isolate each individual weight within the AIMC tile, and then reveal the isolated weight’s value by ADC power side-channel analysis using the Transformer NN model. We demonstrate the practical applicability and robustness of our proposed method by power side-channel analysis of Oscillator-based ADCs that are typically integrated within AIMC tiles. The experimental results show high accuracy and robustness of our Transformer-based analysis, implying potential vulnerabilities in AIMC tiles. Roozbeh Siyadatzadeh, Fatemeh Mehrafrooz, Nele Mentens, Todor P. Stefanov |
DSD | 3 |
| 2025 | SPARK: Secure Privacy-Preserving Anonymous Swarm Attestation for In-Vehicle NetworksabstractIn recent years, vehicles have evolved into cyberphysical autonomous systems that rely on sensor data from various sources within the vehicle. With the emergence of Vehicle-to-Everything (V2X) technology, the scope of the collaborative functionality in vehicles is now expanding to the inter-vehicular level. To support these modern capabilities, the complexity of the Electronic Control Units (ECUs) and the In-Vehicle Network (IVN) architecture is rapidly increasing. As a result, IVNs are now swarms of devices that communicate safety-critical data. Unfortunately, current vehicular networks lack security, opening the path to numerous cyberattacks. A typical solution for verifying the integrity of multiple devices is swarm attestation. However, in a typical IVN setting, only the Original Equipment Manufacturer (OEM) has access to the legitimate configuration of the ECUs and does not want to disclose this information due to intellectual property and security concerns. Therefore, state- of-the-art swarm attestation schemes, which do not provide privacy guarantees, are unsuitable for IVNs.This paper proposes Secure Privacy Preserving Anonymous Swarm Attestation for In-Vehicle Networks (SPARK), which builds upon a novel group signature scheme to enable privacy-preserving, anonymous, and traceable swarm attestation of IVNs. We validate SPARK through a proof-of-concept implementation using a standardized hardware Trusted Platform Module (TPM 2.0) and representative hardware platforms. The results demonstrate the real-world applicability of SPARK. Wouter Hellemans, Nada El Kassem, Md Masoom Rabbani, Edlira Dushku, Liqun Chen 0002, An Braeken, Bart Preneel, Nele Mentens |
EuroS&P | 8 |
| 2025 | Designing Hardware-Friendly Hash Functions for Network Security Using Cartesian Genetic Programming
Jo Vliegen, Stjepan Picek, Nele Mentens |
EvoApplications (2) | 4 |
| 2025 | ITERATOR: Interruptible Remote Attestation Through Cuckoo FiltersabstractRemote attestation (RA) is emerging as a promising security mechanism that establishes trust in IoT devices by detecting the malware presence. Typically, RA consists of computing a hash over the device’s memory and is executed as anatomicprocedure to guarantee the reliability of the attestation evidence. However, in real-world situations, such as those involving real-time systems, energy-harvesting devices, or mission-critical operations, the IoT device may not be able to complete the attestation procedure due to various factors like task scheduling, limited battery life, or higher priority tasks. In such scenarios where flexibility, adaptability, and security are paramount, enablinginterruptibilityof RA is crucial. This paper presents a novel approach called ITERATOR which leverages hash-based storage to enable interruptible RA without any additional hardware requirements. Our proposal transforms the device attestation procedure from the traditional approach of memory hash computation to a lookup operation in a hash-based storage, namely, Cuckoo filter. The ITERATOR protocol divides the device’s memory into blocks associated with a Cuckoo filter bucket. This approach allows the device to perform RA in multiple rounds, ensuring secure interruptible attestation. We perform software simulations of ITERATOR, demonstrating its high effectiveness in detecting the malware presence. Due to its interruptible design, ITERATOR cannot guarantee 100% detection in a single attestation round; however, repeated rounds make long-term evasion by malware highly unlikely. In particular, the experiments showed that the probability of evading the detection ranges between 37% and less than 1%, depending on the protocol configuration. Moreover, we validate ITERATOR’s efficiency through two hardware proof-of-concept implementations that rely on ESP32 and FPGA platforms. The FPGA implementation shows the high efficiency of the protocol, with 34.3ns to attest a single memory block. Nicoló Sponziello, Arish Sateesan, Md Masoom Rabbani, Nele Mentens, Nicola Dragoni, Edlira Dushku |
IEEE Internet Things J. | 4 |
| 2025 | Toward a Real-Time Intrusion Detection System for Modern In-Vehicle NetworksabstractOver the past decade, it has been demonstrated that the In-Vehicle Network (IVN) of a modern Intelligent Transportation System (ITS) is vulnerable to several cyberattacks. Given the collaborative nature of these systems, detecting (remote) cyberattacks is of utmost importance in ensuring trusted interactions. One key technique that has been explored to detect adversarial presence in IVNs are Intrusion Detection Systems (IDSs). However, many existing solutions focus on legacy architectures or are not practically feasible due to their hardware requirements or inability to operate in real-time. To this end, we propose Modular Reduced Temporal Convolutional Network (MR-TCN), an efficient IDS architecture that can effectively be accelerated on hardware to enable real-time intrusion detection in low-cost embedded platforms. Additionally, we evaluate variants of MR-TCN on a Field-Programmable Gate Array (FPGA) platform across a diverse range of IVN traffic (i.e., CAN CC, CAN FD, and Automotive Ethernet), demonstrating its suitability in real-world applications. Wouter Hellemans, Laurens Le Jeune, Md Masoom Rabbani, Bart Preneel, Nele Mentens |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Optimised AES with RISC-V Vector ExtensionsabstractWith the advent of quantum computers, organizations and users should consider the potential impact of quantum threats on their cryptographic systems and be prepared to adopt Post-Quantum Cryptography (PQC) solutions when needed. However, PQC algorithms are often difficult to implement on standard processors and resource-constrained embedded devices, due to complicated mathematical algorithms and large parameters. The goal of this research is to design efficient HW/SW co-design of the PQC algorithm Classic McEliece (CM) using the RISC-V Instruction Set Architecture (ISA). In the first step, the acceleration of the AES algorithm, which is used as part of the key generation in CM, is explored using RISC-V Vector Extensions version 1.0 (RVV1.0). In this paper, we compare the vector-accelerated AES running on Vicuan coprocessor with the scalar AES running on Ibex. Mahnaz Namazi Rizi, Nusa Zidaric, Lejla Batina, Nele Mentens |
DDECS | 4 |
| 2024 | A Systematic Exploration of Evolutionary Computation for the Design of Hardware-oriented Non-cryptographic Hash FunctionsabstractNon-cryptographic (NC) hash functions are crucial in high-speed search applications and probabilistic data structures (PDS) such as Bloom filters and Count-Min sketches for efficient lookups and counting. These operations necessitate execution at line rates to accommodate the high-speed demands of Terabit Ethernet networks, characterized by bandwidths exceeding 100 Gbps. Consequently, a growing inclination towards hardware platforms, particularly Field Programmable Gate Arrays (FPGAs), is evident in network security applications. Given the centrality of hash functions in these structures, any enhancements to their design carry substantial implications for overall system performance. However, hash functions must exhibit independence, uniform distribution, and hardware-friendly characteristics. In this work, we employ Genetic Programming (GP) with avalanche metrics as a fitness function to devise a hardware-friendly family of NC hash functions called the Evolutionary hash (E-hash). We provide a detailed experimental analysis to offer insights on primitive set combinations involving logical operations and diverse hyperparameter settings, encompassing variables such as the number of nodes, tree height, population size, crossover and mutation rate, tournament size, number of constants, and generations. Compared to existing state-of-the-art hardware-friendly hash functions, the proposed E-hash family exhibits an 8.4% improvement in terms of operating frequency and throughput and 7.74% in latency on FPGA. Jo Vliegen, Stjepan Picek, Nele Mentens |
GECCO | 4 |
| 2024 | SPArch: A Hardware-oriented Sketch-based Architecture for High-speed Network Flow MeasurementsabstractNetwork flow measurement is an integral part of modern high-speed applications for network security and data-stream processing. However, processing at line rate while maintaining the required data structure within the on-chip memory of the hardware platform is a challenging task for measurement algorithms, especially when accuracy is of primary importance, such as in network security applications. Most of the existing measurement algorithms are no exception to such issues when deployed in high-speed networking environments and are also not tailored for efficient hardware implementation. Sketch-based measurement algorithms minimize the memory requirement and are suitable for high-speed networks but possess a low memory-accuracy trade-off and lack the versatility of individual flow mapping. To address these challenges, we present a hardware-friendly data structure named Sketch-based Pseudo-associative array Architecture (SPArch). SPArch is highly accurate and extremely memory-efficient, making it suitable for network flow measurement and security applications. The parallelism in SPArch ensures minimal and constant memory access cycles. Unlike other sketch architectures, SPArch provides the functionality of individual flow mapping similar to associative arrays, and the optimized version of SPArch allows the organization of counters in multiple buckets based on the flow sizes. An in-depth analysis of SPArch is carried out in this article and implemented SPArch on the Alveo data center accelerator card, demonstrating its suitability for high-speed networks. Arish Sateesan, Jo Vliegen, Simon Scherrer, Hsu-Chun Hsiao, Adrian Perrig, Nele Mentens |
ACM Trans. Priv. Secur. | 6 |
| 2024 | Introduction to the Special Issue on FPGA-based Embedded Systems for Industrial and IoT ApplicationsabstractNo abstract available. Satwant Singh, Carlos Enrique Montenegro-Marín, Yun Liang 0001, Yao Chen 0008, Nele Mentens, Raymond X. Nijssen |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2023 | Preface ASAP 2023abstractkeynote speeches and provide information regarding the ASAP'2023 organizing, steering and program committees, subreviewers, and sponsors. João M. P. Cardoso, Alexandra Jimborean, Nele Mentens, José Gabriel F. Coutinho |
ASAP | 3 |
| 2023 | Yes we CAN!: Towards bringing security to legacy-restricted Controller Area Networks. A reviewabstractWith the demand for advanced functionality such as autonomous driving, the complexity and connectivity of modern vehicles have faced an overwhelming expansion in recent years. Although the numerous interfaces pave the way for a better user experience, recent research has demonstrated that they can also serve as an attack surface for cybercriminals. Therefore, researchers have been challenged to develop a wide variety of security solutions aiming to solve specific issues. Wouter Hellemans, Md Masoom Rabbani, Bart Preneel, Nele Mentens |
CF | 4 |
| 2023 | NimbleAI: Towards Neuromorphic Sensing-Processing 3D-integrated ChipsabstractThe NimbleAI Horizon Europe project leverages key principles of energy-efficient visual sensing and processing in biological eyes and brains, and harnesses the latest advances in$\mathbf{33D}$stacked silicon integration, to create an integral sensing-processing neuromorphic architecture that efficiently and accurately runs computer vision algorithms in area-constrained endpoint chips. The rationale behind the NimbleAI architecture is: sense data only with high information value and discard data as soon as they are found not to be useful for the application (in a given context). The NimbleAI sensing-processing architecture is to be specialized after-deployment by tunning system-level trade-offs for each particular computer vision algorithm and deployment environment. The objectives of NimbleAI are: (1)$\mathbf{100x}$performance per mW gains compared to state-of-the-practice solutions (i.e., CPU/GPUs processing frame-based video); (2)$\mathbf{50x}$processing latency reduction compared to CPU/GPUs; (3) energy consumption in the order of tens of mWs; and (4) silicon area of approx. 50 mm2. Xabier Iturbe, Nassim Abderrahmane, Jaume Abella 0001, Sergi Alcaide, Eric Beyne, Henri-Pierre Charles, Christelle Charpin-Nicolle, Lars Chittka, Angélica Dávila, Arne Erdmann, Carles Estrada, Ander Fernández, Anna Fontanelli, José Flich, Gianluca Furano, Alejandro Hernán Gloriani, Erik Isusquiza, Radu Grosu, Carles Hernández 0001, Daniele Ielmini, Maha Kooli, Nicola Lepri, Bernabé Linares-Barranco, Jean-Loup Lachese, Eric Laurent, Menno Lindwer, Frank Linsenmaier, Mikel Luján, Karel Masarík, Nele Mentens, Orlando Moreira, Chinmay Nawghane, Luca Peres, Jean-Philippe Noël, Arash Pourtaherian, Christoph Posch, Peter Priller, Zdenek Prikryl, Felix Resch, Oliver Rhodes, Todor P. Stefanov, Moritz Storring, Michele Taliercio, Rafael Tornero, Marcel D. van de Burgwal, Geert Van der Plas, Elisa Vianello, Pavel Zaykov |
DATE | 31 |
| 2023 | Maximizing the Potential of Custom RISC-V Vector Extensions for Speeding up SHA-3 Hash FunctionsabstractSHA-3 is considered to be one of the most secure standardized hash functions. It relies on the Keccak-f[1 600] permutation, which operates on an internal state of 1 600 bits, mostly represented as a 5 x 5 x 64-bit matrix. While existing implementations process the state sequentially in chunks of typically 32 or 64 bits, the Keccak-f[1 600] permutation can benefit a lot from speedup through parallelization. This paper is the first to explore the full potential of parallelization of Keccak-f[1 600] in RISC-V based processors through custom vector extensions on 32-bit and 64-bit architectures. We analyze the Keccak$\mathbf{f}[1 \ 600]$permutation, composed of five different step mappings, and propose ten custom vector instructions to speed up the computation. We realize these extensions in a SIMD processor described in System Verilog. We compare the performance of our designs to existing architectures based on vectorized application-specific instruction set processors (ASIP). We show that our designs outperform all related work in throughput due to our carefully selected custom vector instructions. Huimin Li 0004, Nele Mentens, Stjepan Picek |
DATE | 2 |
| 2023 | A Survey of Recent Developments in Testability, Safety and Security of RISC-V ProcessorsabstractWith the continued success of the open RISC-V architecture, practical deployment of RISC-V processors necessitates an in-depth consideration of their testability, safety and security aspects. This survey provides an overview of recent developments in this quickly-evolving field. We start with discussing the application of state-of-the-art functional and system-level test solutions to RISC-V processors. Then, we discuss the use of RISC-V processors for safety-related applications; to this end, we outline the essential techniques necessary to obtain safety both in the functional and in the timing domain and review recent processor designs with safety features. Finally, we survey the different aspects of security with respect to RISC-V implementations and discuss the relationship between cryptographic protocols and primitives on the one hand and the RISC-V processor architecture and hardware implementation on the other. We also comment on the role of a RISC-V processor for system security and its resilience against side-channel attacks. Jens Anders, Pablo Andreu, Bernd Becker 0001, Steffen Becker 0001, Riccardo Cantoro, Nikolaos Ioannis Deligiannis, Nourhan Elhamawy, Tobias Faller, Carles Hernández 0001, Nele Mentens, Mahnaz Namazi Rizi, Ilia Polian, Abolfazl Sajadi, Matthias Sauer 0002, Denis Schwachhofer, Matteo Sonza Reorda, Todor Stefanov, Ilya Tuzov, Stefan Wagner 0001, Nusa Zidaric |
ETS | 10 |
| 2023 | Evolving Non-cryptographic Hash Functions Using Genetic Programming for High-speed Lookups in Network Security Applications
Arish Sateesan, Jo Vliegen, Stjepan Picek, Nele Mentens |
EvoApplications@EvoStar | 5 |
| 2023 | PrefaceabstractThis book contains the proceedings of the 33rd edition of the International Conference on Field Programmable Logic and Applications (FPL) held in Gothenburg on September 4th to 8th, 2023. Ioannis Sourdis, Nele Mentens, Leonel Sousa, Pedro Trancoso |
FPL | 2 |
| 2023 | ALBUS: a Probabilistic Monitoring Algorithm to Counter Burst-Flood AttacksabstractModern DDoS defense systems rely on probabilistic monitoring algorithms to identify flows that exceed a volume threshold and should thus be penalized. Commonly, classic sketch algorithms are considered sufficiently accurate for usage in DDoS defense. However, as we show in this paper, these algorithms achieve poor detection accuracy under burst-flood attacks, i.e., volumetric DDoS attacks composed of a swarm of medium-rate sub-second traffic bursts. Under this challenging attack pattern, traditional sketch algorithms can only detect a high share of the attack bursts by incurring a large number of false positives. In this paper, we present ALBUS, a probabilistic monitoring algorithm that overcomes the inherent limitations of previous schemes: ALBUS is highly effective at detecting large bursts while reporting no legitimate flows, and therefore improves on prior work regarding both recall and precision. Besides improving accuracy, ALBUS scales to high traffic rates, which we demonstrate with an FPGA implementation, and is suitable for programmable switches, which we showcase with a P4 implementation. Simon Scherrer, Jo Vliegen, Arish Sateesan, Hsu-Chun Hsiao, Nele Mentens, Adrian Perrig |
SRDS | 5 |
| 2023 | PROVE: Provable remote attestation for public verifiabilityabstractThe expanding attack surface of Internet of Things (IoT) systems calls for innovative security approaches to verify the reliability of IoT devices. To this end, Remote Attestation (RA) serves as a key mechanism that remotely detects the presence of malware in IoT devices. Typically, RA allows a centralized trusted Verifier to retrieve reliable evidence about the software integrity of an untrusted Prover. Existing RA schemes generally rely on the assumption that the Verifier and the Prover know each other and have pre-shared cryptographic keys during the bootstrap phase. However, these assumptions are not realistic to employ over commonly used event-driven IoT networks, in which the interacting parties do not know each other and do not communicate directly. This paper proposes PROVE, a novel protocol that allows many Verifiers to attest one or more Provers without pre-shared key material and without using public-key cryptography which is often not suitable for resource-constraint IoT devices. In particular, PROVE considers a realistic IoT system where devices adopt the publish/subscribe communication paradigm. In PROVE, the subscribers act as untrusted Verifiers and attest not only the firmware integrity of the publishers that act as untrusted Provers but also the authenticity of the received data originated from these publishers. We simulate PROVE on the Contiki emulator and demonstrate the scalability of the solution. We also validate PROVE through two hardware proof-of-concept implementations: PROVE and PROVE+, which rely on different cryptographic cores. The results show that a complete execution of the protocol takes 4605 ns and 324 ns for PROVE and PROVE+, respectively. Edlira Dushku, Md Masoom Rabbani, Jo Vliegen, An Braeken, Nele Mentens |
J. Inf. Secur. Appl. | 5 |
| 2022 | Energy and side-channel security evaluation of near-threshold cryptographic circuits in 28nm FD-SOI technologyabstractThis paper is the first to present an implementation of a cryptographic circuit in 28nm FD-SOI using near-threshold design. The implemented cipher, Ketje Jr, is a lightweight authenticated encryption algorithm. The energy consumption of representative authenticated encryption operations as well as the information leakage through the power consumption side-channel are evaluated. The results show that an ultra-low energy implementation can be achieved, and that the near-threshold design has little influence on the Signal to Noise Ratio in the power measurements of our chip. Arthur Beckers, Roel Uytterhoeven, Thomas Vandenabeele, Jo Vliegen, Lennert Wouters, Joan Daemen, Wim Dehaene, Benedikt Gierlichs, Nele Mentens |
CF | 9 |
| 2022 | A scalable SIMD RISC-V based processor with customized vector extensions for CRYSTALS-kyberabstractThis paper uses RISC-V vector extensions to speed up lattice-based operations in architectures based on HW/SW co-design. We analyze the structure of the number-theoretic transform (NTT), inverse NTT (INTT), and coefficient-wise multiplication (CWM) in CRYSTALS-Kyber, a lattice-based key encapsulation mechanism. We propose 12 vector extensions for CRYSTALS-Kyber multiplication and four for finite field operations in combination with two optimizations of the HW/SW interface. This results in a speed-up of 141.7, 168.7, and 245.5 times for NTT, INTT, and CWM, respectively, compared with the baseline implementation, and a speed-up of over four times compared with the state-of-the-art HW/SW co-design using RV32IMC. Huimin Li 0004, Nele Mentens, Stjepan Picek |
DAC | 2 |
| 2022 | Feature dimensionality in CNN acceleration for high-throughput network intrusion detectionabstractWith the ever increasing need for better cybersecurity, and due to the continuous growth of network traffic bandwidths, there is a continuous pursuit of faster and smarter network intrusion detection systems. Neural network-based solutions on FPGAs are very effective in detecting different types of attacks, but have problems with analyzing network traffic online at line speed. One important bottleneck that limits the throughput in raw traffic-based existing systems, is the input shape of the features that are extracted from the raw data. In this work, we propose new methods for extracting and representing features based on raw network traffic in online network intrusion detection systems. We show that feature dimensionality has a significant influence on the classification accuracy and the throughput. Our experiments are based on FPGA-based neural networks accelerated through FINN. We compare three newly proposed input shapes to the traditional 2D-based approach, and we show that two of the presented techniques greatly surpass the state-of-the-art with regards to accuracy and throughput. Our best architecture reaches a maximum bandwidth of 23.09 Gbps, while maintaining over 99% accuracy on both the UNSW-NB15 and CICIDS2017 datasets. Laurens Le Jeune, Toon Goedemé, Nele Mentens |
FPL | 3 |
| 2022 | FOCUS: Frequency Based Detection of Covert Ultrasonic Signals
Wouter Hellemans, Md Masoom Rabbani, Jo Vliegen, Nele Mentens |
SEC | 4 |
| 2022 | Introduction to the Special Section on FPL 2020abstractNo abstract available. Nele Mentens, Leonel Sousa, Pedro Trancoso |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2021 | Invited: Security Beyond Bulk Silicon: Opportunities and Challenges of Emerging DevicesabstractWhile traditional chips in bulk silicon technology are widely used for reliable and highly efficient systems, there are applications that call for devices in other technologies. On the one hand, novel device technologies need to be re-evaluated with respect to potential threats and attacks, and how these can be faced with existing and novel security solutions and methods. On the other hand, emerging device technologies bring opportunities for building the secure systems of the future. In this paper, we will give an overview of applications and security primitives developed in three important emerging device technologies, namely memristors, fully depleted silicon on insulator (FD-SOI) and flexible electronics. Lejla Batina, Rosario Cammarota, Nele Mentens, Ahmad-Reza Sadeghi, Martha Johanna Sepúlveda, Shaza Zeitouni |
DAC | 3 |
| 2021 | Novel Non-cryptographic Hash Functions for Networking and Security Applications on FPGAabstractThis paper proposes the design and FPGA implementation of five novel non-cryptographic hash functions, that are suitable to be used in networking and security applications that require fast lookup and/or counting architectures. Our approach is inspired by the design of the existing non-cryptographic hash function Xoodoo-NC, which is constructed through the concatenation of several Xoodoo permutations. We similarly construct non-cryptographic hash functions based on the concatenation of several rounds of symmetric-key ciphers. The goal is to achieve high performance in combination with good avalanche properties, which are required in order to have a significant change in the output value as a result of a limited change in the input value. We simulate how many rounds are needed to achieve satisfactory avalanche scores and we implement the corresponding non-cryptographic hash functions on an FPGA to evaluate the occupied resources and the performance. One of the proposed non-cryptographic hash functions, namely GIFT-NC, outperforms all previously proposed non-cryptographic hash functions in terms of throughput and latency, in exchange for an acceptable increase in FPGA resources. Thomas Claesen, Arish Sateesan, Jo Vliegen, Nele Mentens |
DSD | 4 |
| 2021 | Trusted Configuration in Cloud FPGAsabstractIn this paper we tackle the open paradoxical challenge of FPGA-accelerated cloud computing: On one hand, clients aim to secure their Intellectual Property (IP) by encrypting their configuration bitstreams prior to uploading them to the cloud. On the other hand, cloud service providers disallow the use of encrypted bitstreams to mitigate rogue configurations from damaging or disabling the FPGA. Instead, cloud providers require a verifiable check on the hardware design that is intended to run on a cloud FPGA at the netlist-level before generating the bitstream and loading it onto the FPGA, therefore, contradicting the IP protection requirement of clients. Currently, there exist no practical solution that can adequately address this challenge.We present the first practical solution that, under reasonable trust assumptions, satisfies the IP protection requirement of the client and provides a bitstream sanity check to the cloud provider. Our proof-of-concept implementation uses existing tools and commodity hardware. It is based on a trusted FPGA shell that utilizes less than 1% of the FPGA resources on a Xilinx VCU118 evaluation board, and an Intel SGX machine running the design checks on the client bitstream. Shaza Zeitouni, Jo Vliegen, Tommaso Frassetto, Dirk Koch, Ahmad-Reza Sadeghi, Nele Mentens |
FCCM | 6 |
| 2021 | AITIA: Embedded AI Techniques for Industrial ApplicationsabstractMotivated by an increasing interest from startups in embedded Artificial Intelligence (AI) and by their limited expertise, the AITIA Project targets the development of embedded AI techniques for industrial applications. This extended abstract presents the motivation and the solutions being developed towards four use cases: smart sensors, network intrusion detection, driver-assistance systems, and Industry 4.0. Marcelo Brandalero, Mitko Veleski, Hector Gerardo Muñoz Hernandez, Muhammad Ali 0010, Laurens Le Jeune, Toon Goedemé, Nele Mentens, Jurgen Vandendriessche, Lancelot Lhoest, Bruno da Silva 0001, Abdellah Touhafi, Diana Göhringer, Michael Hübner 0001 |
FPL | 7 |
| 2021 | Speed Records in Network Flow Measurement on FPGAabstractNetwork traffic measurement keeps track of the amount of traffic sent by each flow in the network. It is a core functionality in applications such as traffic engineering and network intrusion detection. In high-speed networks, it is impossible to keep an exact count of the flow traffic, due to limitations with respect to memory and computational speed. Therefore, probabilistic data structures, such as sketches, are used. This paper proposes Approximate Count-Min sketch or ACM sketch, a novel variant of the Count-Min sketch algorithm that uses less memory and has a higher throughput compared to other FPGA-based sketch implementations. A-CM sketch relies on optimizations at two levels: (1) it uses approximate counters and the newly proposed Hardware-oriented Simple Active Counter algorithm to efficiently implement these counters; (2) it uses a distribution of the embedded memory, optimized towards maximum operating frequency. To the best of our knowledge, A-CM sketch outperforms all other FPGA-based sketch implementations. Arish Sateesan, Jo Vliegen, Simon Scherrer, Hsu-Chun Hsiao, Adrian Perrig, Nele Mentens |
FPL | 6 |
| 2021 | RESERVE: Remote Attestation of Intermittent IoT devicesabstractInternet of Things (IoT) devices have enveloped our surroundings and have been increasingly deployed in many domains. Even though the IoT has generated unprecedented opportunities, the poorly secured design of IoT devices makes them an easy target for cyber attacks. Aimed at securing IoT devices, Remote Attestation (RA) is a security technique that identifies threat presence in IoT systems. Typically, RA is an atomic procedure that requires uninterrupted connectivity to execute. However, in energy harvesting context where intermittent IoT devices go into sleep mode immediately after regular operations, the atomic property is difficult to achieve. In this paper, we propose RESERVE, a novel lightweight RA protocol designed specifically for Intermittent IoT devices. RESERVE aims to improve the security of intermittent systems by detecting malware presence during online mode and guaranteeing with some probability software legitimacy during offline mode. In particular, RESERVE ensures trustworthiness by organizing the device's software into modules, and after regular operation each device attests as many modules as fit in its energy budget. Md Masoom Rabbani, Edlira Dushku, Jo Vliegen, An Braeken, Nicola Dragoni, Nele Mentens |
SenSys | 6 |
| 2021 | Low-Rate Overuse Flow Tracer (LOFT): An Efficient and Scalable Algorithm for Detecting Overuse FlowsabstractCurrent probabilistic flow-size monitoring can only detect heavy hitters (e.g., flows utilizing 10 times their permitted bandwidth), but cannot detect smaller overuse (e.g., flows utilizing 50-100 % more than their permitted bandwidth). Thus, these systems lack accuracy in the challenging environment of high-throughput packet processing, where fast-memory resources are scarce. Nevertheless, many applications rely on accurate flow-size estimation, e.g., for network monitoring, anomaly detection and Quality of Service. We design, analyze, implement, and evaluate LOFT, a new approach for efficiently detecting overuse flows that achieves dramatically better properties than prior work. LOFT can detect 1.50x overuse flows in one second, whereas prior approaches can only reliably detect flows that overuse their allocation by at least 3x. We demonstrate LOFT's suitability for high-speed packet processing with implementations in the DPDK framework and on an FPGA. Simon Scherrer, Che-Yu Wu, Yu-Hsi Chiang, Benjamin Rothenberger, Daniele Enrico Asoni, Arish Sateesan, Jo Vliegen, Nele Mentens, Hsu-Chun Hsiao, Adrian Perrig |
SRDS | 8 |
| 2021 | Evaluating the ROCKY Countermeasure for Side-Channel LeakageabstractROCKY is a recently introduced countermeasure against fault attacks for authenticated encryption algorithms. It is based on the random rotation of the internal state. In this work, we evaluate the effectiveness of ROCKY as a countermeasure against side-channel attacks. We implement four different types of FPGA-oriented architectures of Xoodoo: an unprotected version and three different versions protected with ROCKY. Xoodoo is used as round function of Xoodyak, which is a scheme in the NIST lightweight cryptography standardization competition. For the experimental setup, the SAKURA-G target board with Spartan-6 FPGA is used. The evaluation of the results is done through test vector leakage assessment (TVLA). This is the first work looking into the side-channel security of the ROCKY countermeasure. Konstantina Miteloudi, Lukasz Chmielewski, Lejla Batina, Nele Mentens |
VLSI-SoC | 4 |
| 2020 | INVITED: AI Utopia or Dystopia - On Securing AI PlatformsabstractToday we are witnessing the widespread deployment of AI algorithms on many computing platforms already to provide various services, thus driving the growing market for AI-based platforms. On the one end, AI support is demanded for resource-constrained embedded devices, e.g., integrated into smart homes and vehicles. On the other end, hi-tech giants and cloud services require AI platforms with increasing computational power to feed their data-hungry neural networks. Neglecting security and privacy aspects on both such low-end and high-end AI platforms can have devastating consequences for end users (privacy and safety) as well as for the AI service providers (IP theft). The utopia of a world where intelligent devices ease the human life can easily turn into a dystopia where the ownership of personal data is threatened.In recent years, tremendous effort has been invested in the development of security architectures that protect sensitive services in isolated execution contexts, called enclaves, which provide protection beyond that of commodity operating systems. In this paper, we elaborate on the most well-known enclave-based security architectures to protect AI services. We point out their shortcomings in providing the security guarantees needed for existing and emerging AI services and discuss new ideas and research directions. Ghada Dessouky, Patrick Jauernig, Nele Mentens, Ahmad-Reza Sadeghi, Emmanuel Stapf |
DAC | 3 |
| 2020 | Breaking a fully Balanced ASIC Coprocessor Implementing Complete Addition Formulas on Weierstrass Elliptic CurvesabstractIn this paper we report on the results of selected horizontal SCA attacks against two open-source designs that implement hardware accelerators for elliptic curve cryptography. Both designs use the complete addition formula to make the point addition and point doubling operations indistinguishable. One of the designs uses in addition means to randomize the operation sequence as a countermeasure. We used the comparison to the mean and an automated SPA to attack both designs. Despite all these countermeasures, we were able to extract the keys processed with a correctness of 100%. Ievgen Kabin, Zoya Dyka, Dan Klann, Nele Mentens, Lejla Batina, Peter Langendörfer |
DSD | 4 |
| 2020 | Novel Bloom filter algorithms and architectures for ultra-high-speed network security applicationsabstractThis paper proposes novel Bloom filter algorithms and FPGA architectures for high-speed searching applications. A Bloom filter is a memory structure that is used to test whether input search data are present in a table of stored data. Bloom filters are extensively used in network security solutions that apply traffic flow monitoring or deep packet inspection. Improving the speed of Bloom filters can therefore have a significant impact on the speed of many network applications. The most important components determining the speed of Bloom filters are hash functions. While hash functions in Bloom filters do not require strong cryptographic properties, they do need a minimized computational delay. We take on the challenge of developing ultra-high-speed Bloom filters on FPGAs by proposing a new noncryptographic hash function, called Xoodoo-NC, derived from the cryptographic permutation Xoodoo. Xoodoo-NC is a reducedround, reduced-state version of Xoodoo, inheriting Xoodoo's desired avalanche properties and low logical depth, resulting in an ultra-low-latency non-cryptographic hash function. We evaluate the performance of Bloom filter architectures based on Xoodoo-NC on a Xilinx UltraScale+FPGA and we compare the performance and resource occupation to existing Bloom filter implementations. We additionally compare our results to memories that use the built-in CAM cores in Xilinx UltraScale+ FPGAs. Our proposed algorithmic and architectural advances lead to Bloom filters that, to the best of our knowledge, outperform all other FPGA-based solutions. Arish Sateesan, Jo Vliegen, Joan Daemen, Nele Mentens |
DSD | 4 |
| 2020 | SHeFU: Secure Hardware-Enabled Protocol for Firmware UpdatesabstractFirmware updates are often termed as a panacea to vulnerable Internet-of-Things (IoT) networks, as firmware updates can fix the exposed bugs and prevent them from being exploited in the future. However, a secure firmware update is a challenging task as IoT devices are often employed in unattended networks. Moreover, malicious updates of firmware in any of the devices of a network, or the non-execution of an update, can create havoc. Although security mechanisms like remote attestation (RA) are quite popular to identify malicious nodes in a network, they are costly in terms of computation/memory usage and communication overhead. To overcome these issues, we propose a “Secure Hardware-enabled Protocol for Firmware Updates (SHeFU)”. The aim of the proposed protocol is two-fold: 1) we obviate the need for remote attestation, and 2) we make sure that malicious nodes are isolated from benign nodes. Assuming a restricted threat model and network constellation, SHeFU ensures secure firmware updates and prevents compromised nodes from communicating with benign nodes in a network. Md Masoom Rabbani, Jo Vliegen, Mauro Conti, Nele Mentens |
ISCAS | 4 |
| 2020 | Lightweight Ciphers and Their Side-Channel ResilienceabstractSide-channel attacks represent a powerful category of attacks against cryptographic devices. Still, side-channel analysis for lightweight ciphers is much less investigated than for instance for AES. Although intuition may lead to the conclusion that lightweight ciphers are weaker in terms of side-channel resistance, that remains to be confirmed and quantified. In this paper, we consider various side-channel analysis metrics which should provide an insight on the resistance of lightweight ciphers against side-channel attacks. In particular, for the non-profiled scenario we use the theoretical confusion coefficient and empirical optimal distinguisher. Our study considers side-channel attacks on the first, the last, or both rounds simultaneously. Furthermore, we conduct a profiled side-channel analysis using various machine learning attacks to recover 4-bit and 8-bit intermediate states of the cipher. Our results show that the difference between AES and lightweight ciphers is smaller than one would expect, and even find scenarios in which lightweight ciphers may be more resistant. Interestingly, we observe that the studied 4-bit S-boxes have a different side-channel resilience, while the difference in the 8-bit ones is only theoretically present. Annelie Heuser, Stjepan Picek, Sylvain Guilley, Nele Mentens |
IEEE Trans. Computers | 4 |
| 2019 | In Hardware We Trust: Gains and Pains of Hardware-assisted SecurityabstractData processing and communication in almost all electronic systems are based on Central Processing Units (CPUs). In order to guarantee confidentiality and integrity of the software running on a CPU, hardware-assisted security architectures are used. However, both the threat model and the non-functional platform requirements, i.e. performance and energy budget, differ when we go from high-end desktop computers and servers to low-end embedded devices that populate the internet of things (IoT). For high-end platforms, a relatively large energy budget is available to protect software against attacks. However, measures to optimize performance give rise to microarchitectural side-channel attacks. IoT devices, in contrast, are constrained in terms of energy consumption and do not incorporate the performance enhancements found in high-end CPUs. Hence, they are less likely to be susceptible to microarchitectural attacks, but give rise to physical attacks, exploiting, e.g., leakage in power consumption or through fault injection. Whereas previous work mostly concentrates on a specific architecture, this paper covers the whole spectrum of computing systems, comparing the corresponding hardware architectures, and most relevant threats. Lejla Batina, Patrick Jauernig, Nele Mentens, Ahmad-Reza Sadeghi, Emmanuel Stapf |
DAC | 3 |
| 2019 | SACHa: Self-Attestation of Configurable HardwareabstractDevice attestation is a procedure to verify whether an embedded device is running the intended application code. This way, protection against both physical attacks and remote attacks on the embedded software is aimed for. With the wide adoption of Field-Programmable Gate Arrays or FPGAs, hardware also became configurable, and hence susceptible to attacks (just like software). In addition, an upcoming trend for hardware-based attestation is the use of configurable FPGA hardware. Therefore, in order to attest a whole system that makes use of FPGAs, the status of both the software and the hardware needs to be verified, without the availability of a tamper-resistant hardware module.In this paper, we propose a solution in which a prover core on the FPGA performs an attestation of the entire FPGA, including a self-attestation. This way, the FPGA can be used as a tamper-resistant hardware module to perform hardware-based attestation of a processor, resulting in a protection of the entire hardware/software system against malicious code updates. Jo Vliegen, Md Masoom Rabbani, Mauro Conti, Nele Mentens |
DATE | 4 |
| 2019 | Dynamic Logic Reconfiguration Based Side-Channel Protection of AES and SerpentabstractDynamic logic reconfiguration is a concept which allows for efficient on-the-fly modifications of combinational circuit behaviour in both ASIC and FPGA devices. The reconfiguration of Boolean functions is achieved by modification of their generators (e.g. shift register-based look-up tables) and it can be controlled from within the chip, without the necessity of any external intervention. This hardware polymorphism can be utilized for the implementation of side-channel attack countermeasures, as demonstrated by Sasdrich et al. for the lightweight cipher PRESENT. In this work we adopt these countermeasures to two of the AES finalists, namely Rijndael and Serpent. Just like PRESENT, both Rijndael and Serpent are block ciphers based on a substitution-permutation network. We describe the countermeasures and adjustments necessary to protect these ciphers using the resources available in modern Xilinx FPGAs. We describe our VHDL implementations and evaluate the side-channel leakage and effectiveness of different countermeasure combinations using a methodology based on Welch's t-test. We did not detect any significant leakage from the fully protected versions of our implementations. We show that the countermeasures proposed by Sasdrich et al. are, with some modifications compared to the protected PRESENT implementation, successfully applicable to AES and Serpent. Petr Socha, Jan Brejník, Stanislav Jerabek, Martin Novotný, Nele Mentens |
DSD | 5 |
| 2019 | SHeLA: Scalable Heterogeneous Layered AttestationabstractThis article proposes a novel mechanism for swarm attestation, i.e., the remote attestation (RA) of a multitude of interconnected devices, also called a swarm of devices. Classical RA protocols work with one prover and one verifier. Swarm attestation protocols assume that the devices in the swarm act both as verifier and prover in order to attest the software integrity of all the devices to a root verifier, typically in a spanning-tree topology. We propose “scalable heterogeneous layered attestation (SHeLA),” a novel RA technique for swarms. Our approach consists of introducing an additional edge layer in between the root verifier and the swarm devices. The edge layer consists of geographically spread devices with a larger computational power and storage capacity than the swarm devices. The main challenges we address are related to the scalability of the swarm, the availability or visibility of the nodes (especially when they are mobile), the heterogeneity of the devices with respect to the wireless communication protocol and interface, and the granularity of the attestation in terms of detecting the sanity of individual swarm devices. We build a proof-of-concept network that allows us to evaluate the computational delay and the resource overhead of the edge and swarm devices, and to perform a thorough security analysis. Md Masoom Rabbani, Jo Vliegen, Jori Winderickx, Mauro Conti, Nele Mentens |
IEEE Internet Things J. | 5 |
| 2018 | HeComm: End-to-end secured communication in a heterogeneous IoT environment via fog computingabstractIn this paper we describe the HeComm (Heterogeneous Communication) architecture. The HeComm architecture utilizes fog computing to provide end-to-end secured communication between IoT nodes residing in networks that operate over different communication protocols. While the fog serves as a bridge between the heterogeneous networks, it does not have access to the confidential data that are being exchanged between the IoT nodes. The HeComm architecture can be divided into two parts: the HeComm protocol and the HeComm communication. The HeComm protocol allows IoT network managers to establish a secret key that is used in the HeComm communication. The HeComm communication secures the exchanged data between the IoT nodes by using object security. This enables an end-to-end protection of the communicated data without putting all trust in the fog. We evaluate the HeComm architecture through a security analysis and via a proof-of-concept implementation that connects an IoT node in a 6LoWPAN network to a node in a LoRaWAN network. The results indicate that the IoT nodes in the HeComm architecture meet the requirements of Class 1 constrained nodes. Furthermore, the HeComm architecture comprises a combination of widely used algorithms and protocols that are often supported in modern IoT devices and platforms, which facilitates the practical deployment of our solution. Jori Winderickx, Dave Singelée, Nele Mentens |
CCNC | 3 |
| 2018 | Digital signatures and signcryption schemes on embedded devices: a trade-off between computation and storageabstractThis paper targets the efficient implementation of digital signatures and signcryption schemes on typical internet-of-things (IoT) devices, i.e. embedded processors with constrained computation power and storage. Both signcryption schemes (providing digital signatures and encryption simultaneously) and digital signatures rely on computation-intensive public-key cryptography. When the number of signatures or encrypted messages the device needs to generate after deployment is limited, a trade-off can be made between performing the entire computation on the embedded device or moving part of the computation to a precomputation phase. The latter results in the storage of the precomputed values in the memory of the processor. We examine this trade-off on a health sensor platform and we additionally apply storage encryption, resulting in five implementation variants of the considered schemes. Jori Winderickx, An Braeken, Dave Singelée, Roel Peeters, Thijs Vandenryt, Ronald Thoelen, Nele Mentens |
CF | 7 |
| 2018 | Design of a Fully Balanced ASIC Coprocessor Implementing Complete Addition Formulas on Weierstrass Elliptic CurvesabstractThis paper discusses the first design of an ASIC coprocessor for Elliptic Curve Cryptography (ECC) using the complete addition law of Renes et al. The main reason for using the complete addition law is the reduced vulnerability to side-channel analysis (SCA) attacks, since point addition and point doubling can be performed with the same addition formulas. Further, all inputs are valid, so there is no need for conditional statements handling special cases such as the point at infinity. The proposed hardware architecture is optimized for area efficiency, targeting applications such as smart cards and RFID tags. A bottom-up design approach is used, minimizing the total implementation area by optimizations in each abstraction layer. The design implements a full-word Montgomery Multiplier ALU (MMALU) with built-in adder functionality. Additionally, an exploration is done on the design parameters of the MMALU and the scheduling of the modular operations in order to minimize the size of the register file. For point multiplication, a Montgomery ladder is implemented with the option of randomizing the execution order of the point operations as a countermeasure against SCA attacks. The post-synthesis implementation results are generated using the open source NANGATE45 library. Niels Pirotte, Jo Vliegen, Lejla Batina, Nele Mentens |
DSD | 4 |
| 2018 | Rethinking Secure FPGAs: Towards a Cryptography-Friendly Configurable Cell Architecture and Its Automated Design FlowabstractThis work proposes the first fine-grained configurable cell array specifically tailored for the implementation of cryptographic algorithms that can be configured using widely adopted hardware description languages. Our solution can be added as a small, crypto-friendly reconfigurable hardware block to be included as an application-specific configurable building block in the next generation of FPGAs, exactly like DSP slices and embedded memory blocks were added in the past. Another application scenario uses our configurable cell array as a small embedded FPGA (eFPGA) which we envision to be added to an ASIC design or a microprocessor. This will solve the need for so-called cryptographic agility, allowing cryptographic algorithms to be upgraded or updated depending on newly detected vulnerabilities or changing standards. We focus on block ciphers and we derive the most suitable cell structure for mapping state-of-the-art algorithms. We develop the related automated design flow, exploiting the synthesis capabilities of Synopsys Design Compiler. We evaluate the performance of our solution by mapping a number of well-known ciphers onto our new cells. The obtained results show that the proposed architecture drastically outperforms commercial FPGAs in terms of silicon area and configuration memory resources, while obtaining a similar throughput. Nele Mentens, Edoardo Charbon, Francesco Regazzoni 0001 |
FCCM | 1 |
| 2017 | Area-optimized montgomery multiplication on IGLOO 2 FPGAsabstractThis paper presents the first area-optimized Montgomery modular multiplication module on low-power reconfigurable IGLOO® 2 FPGAs, from Microsemi. In order to obtain a good response time with few resources, the FPGA pipelined Math blocks and the embedded memory blocks are fully leveraged. As a result, 256-bit modular multiplications can be done in 2.33 μs, at a cost of 505 LUT4 cells, 257 Flip Flops, 1 Math block and 1 64×18 RAM block. If more area resources are considered, a modular multiplication can be performed in 1.25 μ8 at a cost of 680 LUT4s, 341 Flip Flops, 2 Math blocks and 2 64×18 RAM blocks. This work is the first fundamental step towards area-efficient public-key cryptography on the Microsemi IGLOO® 2 FPGAs. Pedro Maat Costa Massolino, Lejla Batina, Ricardo Chaves, Nele Mentens |
FPL | 4 |
| 2017 | The Monte Carlo PUFabstractPhysically unclonable functions are used for IP protection, hardware authentication and supply chain security. While many PUF constructions have been put forward in the past decade, only few of them are applicable to FPGA platforms. Strict constraints on the placement and routing are the main disadvantages of the existing PUFs on FPGAs, because they place a high effort on the designer. In this paper we propose a new delay-based PUF construction called Monte Carlo PUF, that does not require low-level placement and routing control. This construction relies on the on-chip Monte Carlo method that is applied for measuring the delays of logic elements in order to extract a unique device fingerprint. The proposed construction allows a trade-off between the evaluation time and the error rate. The Monte Carlo PUF is implemented and evaluated on Xilinx Spartan-6 FPGAs. Vladimir Rozic, Bohan Yang 0001, Jo Vliegen, Nele Mentens, Ingrid Verbauwhede |
FPL | 4 |
| 2017 | Side-channel analysis and machine learning: A practical perspectiveabstractThe field of side-channel analysis has made significant progress over time. Side-channel analysis is now used in practice in design companies as well as in test laboratories, and the security of products against side-channel attacks has significantly improved. However, there are still some remaining issues to be solved for side-channel analysis to become more effective. Side-channel analysis consists of two steps, commonly referred to as identification and exploitation. The identification consists of understanding the leakage and building suitable models. The exploitation consists of using the identified leakage models to extract the secret key. In scenarios where the model is poorly known, it can be approximated in a profiling phase. There, machine learning techniques are gaining value. In this paper, we conduct extensive analysis of several machine learning techniques, showing the importance of proper parameter tuning and training. In contrast to what is perceived as common knowledge in unrestricted scenarios, we show that some machine learning techniques can significantly outperform template attacks when properly used. We therefore stress that the traditional worst case security assessment of cryptographic implementations, that mainly includes template attacks, might not be accurate enough. Besides that, we present a new measure called the Data Confusion Factor that can be used to assess how well machine learning techniques will perform on a certain dataset. Stjepan Picek, Annelie Heuser, Alan Jovic, Simone A. Ludwig, Sylvain Guilley, Domagoj Jakobovic, Nele Mentens |
IJCNN | 7 |
| 2016 | PRNGs for Masking Applications and Their Mapping to Evolvable Hardware
Stjepan Picek, Bohan Yang 0001, Vladimir Rozic, Jo Vliegen, Jori Winderickx, Thomas De Cnudde, Nele Mentens |
CARDIS | 7 |
| 2016 | TOTAL: TRNG on-the-fly testing for attack detection using Lightweight hardware
Bohan Yang 0001, Vladimir Rozic, Nele Mentens, Wim Dehaene, Ingrid Verbauwhede |
DATE | 3 |
| 2016 | Evolutionary Algorithms for Finding Short Addition Chains: Going the Distance
Stjepan Picek, Carlos A. Coello Coello, Domagoj Jakobovic, Nele Mentens |
EvoCOP | 4 |
| 2016 | Exploring the use of shift register lookup tables for Keccak implementations on Xilinx FPGAsabstractWe explore the possibility of using shift register lookup tables (SRLs) for the implementation of Keccak on Xilinx FPGAs. The approach originates from the observation that the ρ step in combination with the state storage can be implemented as a collection of shift registers. This way, we achieve a slice-wise implementation using 25 shift registers of various lengths, resulting in 75 32-bit and 6 16-bit SRL primitives on a Virtex-5. This approach, however, does not comply efficiently with the common interface of Keccak. We therefore propose to utilize a modified interface in order to avoid the redundant storage of the state. This modified version of Keccak, with equivalent cryptographic strength, can be implemented without Block RAM, outperforming all previously proposed implementations in terms of the number of slices. Furthermore it outperforms several other lightweight slice-wise implementations in terms of throughput. Jori Winderickx, Joan Daemen, Nele Mentens |
FPL | 3 |
| 2016 | Evolving Cryptographic Pseudorandom Number Generators
Stjepan Picek, Dominik Germek, Vladimir Rozic, Bohan Yang 0001, Domagoj Jakobovic, Nele Mentens |
PPSN | 6 |
| 2016 | On the Construction of Hardware-Friendly 4\times 4 and 5\times 5 S-Boxes
Stjepan Picek, Bohan Yang 0001, Vladimir Rozic, Nele Mentens |
SAC | 4 |
| 2016 | A Search Strategy to Optimize the Affine Variant Properties of S-Boxes
Stjepan Picek, Bohan Yang 0001, Nele Mentens |
WAIFI | 3 |
| 2016 | Guest EditorialabstractInternational audience Nele Mentens, Damien Sauveron, José María Sierra, Shiuh-Jeng Wang, Isaac Woungang |
IET Inf. Secur. | 1 |
| 2015 | Embedded HW/SW platform for on-the-fly testing of true random number generators
Bohan Yang 0001, Vladimir Rozic, Nele Mentens, Wim Dehaene, Ingrid Verbauwhede |
DATE | 3 |
| 2015 | Challenges in designing trustworthy cryptographic co-processorsabstractSecurity is becoming ubiquitous in our society. However, the vulnerability of electronic devices that implement the needed cryptographic primitives has become a major issue. This paper starts by presenting a comprehensive overview of the existing attacks to cryptography implementations. Thereafter, the state-of-the-art on some of the most critical aspects of designing cryptographic co-processors are presented. This analysis starts by considering the design of asymmetrical and symmetrical cryptographic primitives, followed by the discussion on the design and online testing of True Random Number Generation. To conclude, techniques for the detection of Hardware Trojans are also discussed. Ricardo Chaves, Giorgio Di Natale, Lejla Batina, Shivam Bhasin, Baris Ege, Apostolos P. Fournaris, Nele Mentens, Stjepan Picek, Francesco Regazzoni 0001, Vladimir Rozic, Nicolas Sklavos 0001, Bohan Yang 0001 |
ISCAS | 7 |
| 2015 | On-the-fly tests for non-ideal true random number generatorsabstractHardware implementations of statistical tests are needed to detect failures and statistical weaknesses of entropy sources in True Random Number Generators on the fly. Current implementations of these tests work under the assumption that the entropy source produces independent, identically distributed (IID) numbers. However, some entropy sources produce non-IID data and rely on compression to provide the full entropy. Currently there are no embedded test implementations suitable for this type of entropy source. We provide the first FPGA implementation of embedded tests that estimate the generated min-Entropy and verify if it is within the expected boundaries. Bohan Yang 0001, Vladimir Rozic, Nele Mentens, Ingrid Verbauwhede |
ISCAS | 3 |
| 2015 | Secure, Remote, Dynamic Reconfiguration of FPGAsabstractWith the widespread availability of broadband Internet, Field-Programmable Gate Arrays (FPGAs) can get remote updates in the field. This provides hardware and software updates, and enables issue solving and upgrade ability without device modification. In order to prevent an attacker from eavesdropping or manipulating the configuration data, security is a necessity. This work describes an architecture that allows the secure, remote reconfiguration of an FPGA. The architecture is partially dynamically reconfigurable and it consists of a static partition that handles the secure communication protocol and a single reconfigurable partition that holds the main application. Our solution distinguishes itself from existing work in two ways: it provides entity authentication and it avoids the use of a trusted third party. The former provides protection against active attackers on the communication channel, while the latter reduces the number of reliable entities. Additionally, this work provides basic countermeasures against simple power-oriented side-channel analysis attacks. The result is an implementation that is optimized toward minimal resource occupation. Because configuration updates occur infrequently, configuration speed is of minor importance with respect to area. A prototype of the proposed design is implemented, using 5,702 slices and having minimal downtime. Jo Vliegen, Nele Mentens, Ingrid Verbauwhede |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | Compact Ring-LWE Cryptoprocessor
Sujoy Sinha Roy, Frederik Vercauteren, Nele Mentens, Donald Donglong Chen, Ingrid Verbauwhede |
CHES | 3 |
| 2013 | Remote FPGA design through eDiViDe - European Digital Virtual Design LababstractThe design and development of digital electronic systems is mainly performed by use of a hardware description language. To prepare students in electrical engineering for a career in hardware design many universities provide courses on VHDL. The traditional approach in teaching VHDL is mainly by means of textbook examples and simulation provided by software applications. These exercises are perceived as monotonous by the students and do not or only very slightly correspond with actual real-life applications based on FPGAs. Moreover, most real-life applications are too expensive to be equipped in student laboratories. To bridge the gap between a simulation-only environment and affordable real-life applications students should be provided access to remote real-life setups with a 24/7 availability and preferably shared between multiple institutes. The eDiViDe platform (European Digital Virtual Design Lab, http://www.edivide.eu), see Fig. 1, provides students with this unlimited and exciting access to FPGA based setups. Instead of theory-only courses and a quick basic lab, they can work their way through digital design courses testing their skills on real-life setups to trigger their interest. The platform hosts multiple FPGA setups at different European institutes. These setups are accessible through a web-based interface with video feedback. VHDL development is performed offline, given an entity and specific setup information. All further steps of the FPGA toolchain are performed on the platform. A reservation system takes care of the FPGA programming and student interaction with the setups. Similar initiatives provide stable solutions with educational support [1,2,3]. The eDiViDe platform differentiates with a distributed platform across several institutes and with the support for advanced setups. It is the result of a joint effort and easily expandable with additional setups at any location. At this moment following setups are available: greenhouse, stepper motor control, sea noise emulator, state machine workshop, Geffe generator, pong / game of life, traffic light control, MIPS CPU. This set will be extended with more advanced setups that include e.g. a partial reconfiguration workshop for audio/video filters, a side-channel analysis setup and a mars rover playfield. Besides promoting digital design education, the eDiViDe platform creates a channel to make the research activities in the contributing universities more visible. Industry could also benefit from this platform to promote their brand and products to soon to be engineers. Jochen Vandorpe, Jo Vliegen, Ruben Smeets, Nele Mentens, Milos Drutarovský, Michal Varchola, Kerstin Lemke-Rust, Paul-Gerhard Plöger, Peter Samarin, Dirk Koch, Yngve Hafting, Jim Tørresen |
FPL | 4 |
| 2012 | Design space exploration for automatically generated cryptographic hardware using functional languagesabstractThis paper presents an EDA (Electronic Design Automation) tool that generates basic building blocks for cryptographic hardware in VHDL. The purpose of the tool is to decrease the design time of cryptographic hardware and to allow designers to make abstraction of both the arithmetic and design complexity. The tool generates multiple implementations for one arithmetic description and then benchmarks the implementations to find the most optimal, based upon design space parameters. These parameters consist of area and speed requirements. We present datapath and control logic results for a Xilinx Virtex-5 FPGA. The novelty in our approach lies in the fact that we exploit the higher-order features of functional languages to facilitate the design space exploration and that we take benefit from the strength of the third-party synthesis tool by generating VHDL code at an abstraction level that is higher than the gate level. Nevertheless, in this stage of the development of the tool, the different cryptographic architectures are hand-made and the selection of the most optimal solution, based upon user requirements, is done by exhaustive search. This means that the tool leaves room for improvement, but forms a solid base for further development. Davy Wolfs, Kris Aerts, Nele Mentens |
FPL | 3 |
| 2010 | A compact FPGA-based architecture for elliptic curve cryptography over prime fieldsabstractThis paper proposes an FPGA-based application-specific elliptic curve processor over a prime field. This research targets applications for which compactness is more important than speed. To obtain a small datapath, the FPGA's dedicated multipliers and carry-chain logic are used and no parallellism is introduced. A small control unit is obtained by following a microcode approach, in which the instructions are stored in the FPGA's Block RAM. The use of algorithms that prevent Simple Power Analysis (SPA) attacks creates an extra cost in latency. Nevertheless, the created processor is flexible in the sense that it can handle all finite field operations over 256-bit prime fields and all elliptic curves of a specified form. The comparison with other implementations on the same generation of FPGAs learns that our design occupies the smallest area. Jo Vliegen, Nele Mentens, Jan Genoe, An Braeken, Serge Kubera, Abdellah Touhafi, Ingrid Verbauwhede |
ASAP | 2 |
| 2009 | Secure FPGA technologies and techniquesabstractThis survey paper proposes an overview of contemporary FPGA-related technologies and techniques that can be used for data and system security. As such we will give an overview of the currently available features in commonly used FPGAs and link these features to established security techniques. The main goal is to evaluate the pros and contras of the different techniques and technologies in order to give directions on the security strategy. An Braeken, Serge Kubera, Frederik Trouillez, Abdellah Touhafi, Nele Mentens, Jo Vliegen |
FPL | 5 |
| 2008 | Power and Fault Analysis Resistance in Hardware through Dynamic Reconfiguration
Nele Mentens, Benedikt Gierlichs, Ingrid Verbauwhede |
CHES | 1 |
| 2007 | Efficient pipelining for modular multiplication architectures in prime fieldsabstractThis paper presents a pipelined architecture of a modular Montgomery multiplier, which is suitable to be used in public key coprocessors. Starting from a baseline implementation of the Montgomery algorithm, a more compact pipelined version is derived. The design makes use of 16-bit integer multiplication blocks that are available on recently manufactured FPGAs. The critical path is optimized by omitting the exact computation of intermediate results in the Montgomery algorithm using a 6-2 carry-save notation. This results in a high-speed architecture,which outperforms previously designed Montgomery multipliers. Because a very popular application of Montgomery multiplication is public key cryptography, we compare our implementation to the state-of-the-art in Montgomery multipliers on the basis of performance results for 1024-bit RSA. Nele Mentens, Kazuo Sakiyama, Bart Preneel, Ingrid Verbauwhede |
ACM Great Lakes Symposium on VLSI | 1 |
| 2007 | Public-Key Cryptography on the Top of a NeedleabstractThis work describes the smallest known hardware implementation for Elliptic/Hyperelliptic Curve Cryptography (ECC/HECC). We propose two solutions for Public-key Cryptography (PKC), which are based on arithmetic on elliptic/hyperelliptic curves. One solution relies on ECC over binary fields 𝔽2𝓃where 𝓃 is a composite number of the form2𝑝(𝑝is a prime) and another on HECC on curves of genus 2 over 𝔽2𝑝. This implies the same arithmetic unit for both cases which supports arithmetic in a field 𝔽2𝑝. Our best solution that still results in a feasible performance features less than 5 kgates with an average power consumption smaller than 10μW. Lejla Batina, Nele Mentens, Kazuo Sakiyama, Bart Preneel, Ingrid Verbauwhede |
ISCAS | 2 |
| 2006 | Fpga-Oriented Secure Data Path Design: Implementation of a Public Key CoprocessorabstractThis paper introduces a secure FPGA implementation of a coprocessor for public key cryptography. It supports Elliptic Curve Cryptography (ECC) as well as the older RSA standard. When choosing adequate key lengths, RSA and ECC are assumed to be secure from an algorithmic point of view. On the other hand, an implementation of these algorithms should also guarantee side-channel security. This feature does not only cause an inevitable performance degradation, but also an area increase. We overcome these drawbacks by fitting the public key architecture and algorithms into a coprocessor that optimally exploites the dedicated features on a Spartan XC3S4000. Although this is a very low-cost FPGA, the performance results of our implementation meet the requirements of a broad range of high-end applications. Nele Mentens, Kazuo Sakiyama, Lejla Batina, Ingrid Verbauwhede, Bart Preneel |
FPL | 1 |
| 2006 | Flexible hardware architectures for curve-based cryptographyabstractThis paper compares implementations of elliptic and hyperelliptic curve cryptography (ECC and HECC) on an FPGA platform. We use the same low-level blocks to implement the basic operations and we choose the bit-lengths so that both systems have equal security levels. The results are in favor of HECC. Our HECC implementation is slightly larger than ECC, but at the same time around 35% faster Lejla Batina, Nele Mentens, Bart Preneel, Ingrid Verbauwhede |
ISCAS | 2 |
| 2005 | Side-channel aware design: Algorithms and Architectures for Elliptic Curve Cryptography over GF(2n)abstractThis paper proposes efficient algorithms for Elliptic Curve Cryptography (ECC). As an example a compact and efficient FPGA architecture for ECC over finite fields of even characteristic is presented. The implementation is balanced in order to increase the security w.r.t. simple side-channel attacks. Multiplication in GF(2 n ), Hardware implementation, Systolic array architecture, Elliptic Curve Cryptography (ECC), Montgomery method for point multiplication © 2005 IEEE. Lejla Batina, Nele Mentens, Bart Preneel, Ingrid Verbauwhede |
ASAP | 2 |
| 2005 | A Systematic Evaluation of Compact Hardware Implementations for the Rijndael S-Box
Nele Mentens, Lejla Batina, Bart Preneel, Ingrid Verbauwhede |
CT-RSA | 1 |
| 2005 | Side-Channel Issues for Designing Secure Hardware ImplementationsabstractSelecting a strong cryptographic algorithm makes no sense if the information leaks out of the device through side-channels. Sensitive information, such as secret keys, can be obtained by observing the power consumption, the electromagnetic radiation, etc. This class of attacks is called side-channel attacks. Another type of attacks, namely fault attacks, reveal secret information by inserting faults into the device. Because both side-channel attacks and fault attacks are based on weaknesses in the implementation, they both belong to the category of implementation attacks. This work gives an overview of the state-of-the-art in implementation attacks, reviews the origin of this problem at the CMOS circuit level and discusses countermeasures. Lejla Batina, Nele Mentens, Ingrid Verbauwhede |
IOLTS | 2 |
| 2004 | An FPGA implementation of an elliptic curve processor GF(2m)abstractThis paper describes a hardware implementation of an arithmetic processor which is efficient for elliptic curve (EC) cryptosystems, which are becoming increasingly popular as an alternative for public key cryptosystems based on factoring. The modular multiplication is implemented using a Montgomery modular multiplication in a systolic array architecture, which has the advantage that the clock frequency becomes independent of the bit length m. Nele Mentens, Siddika Berna Örs Yalçin, Bart Preneel |
ACM Great Lakes Symposium on VLSI | 1 |