EDBT 2026 Demo / reviewers in the wild / expert
Ingrid Verbauwhede
dblp:92/16 · also Ingrid M. R. Verbauwhede
· DBLP profile ↗
248ranked-venue papers
12as first author
38since 2021 · last 2026
0000-0002-0879-076XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 144 · 10 first-author · 20 since 2021Security and privacy · 87 · 1 first-author · 16 since 2021Software engineering, systems software and programming languages · 33 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorTheory of computation · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMOOTHIE: (Multi-)scalar Multiplication Optimisations On TFHE
Xander Pottier, Jan-Pieter D'Anvers, Thomas de Ruijter, Ingrid Verbauwhede |
CRYPTO (2) | 4 |
| 2026 | Extending and Accelerating Inner Product Masking with Fault Detection via Instruction Set ExtensionabstractInner product masking is a well-studied masking countermeasure against side-channel attacks. IPM-FD further extends the IPM scheme with fault detection capabilities. However, implementing IPM-FD in software especially on embedded devices results in high computational overhead. Therefore, in this work we perform a detailed analysis of all building blocks for IPM-FD scheme and propose a Masked Processing Unit to accelerate all operations, for example multiplication and IPM-FD specific Homogenization. We can then offload these computational extensive operations with dedicated hardware support. With only 4.05% and 4.01% increase in Look-Up Tables and Flip-Flops (Random Number Generator excluded), respectively, compared with baseline cv32e40p RISC-V core, we can achieve up to 16.55× speed-up factor with optimal configuration. We then practically evaluate the side-channel security via uni- and bivariate Test Vector Leakage Assessment which exhibits no leakage. Finally, we use two different methods to simulate the injected fault and confirm the fault detection capability of up to k−1 faults, with k being the replication factor. Songqiao Cui, Geng Luo, Junhan Bao, Josep Balasch, Ingrid Verbauwhede |
DATE | 5 |
| 2026 | ML-DSA-OSH: An Efficient, Open-Source Hardware Implementation of ML-DSAabstractML-DSA is a post-quantum lattice-based digital signature algorithm (DSA) that the National Institute of Standards and Technology (NIST) recently standardized as FIPS 204. Remarkably, there are only a handful of published hardware designs and no open-source hardware implementations of complete ML-DSA. In this work, we present an efficient open-source hardware (OSH) design of ML-DSA, based on a Dilithium implementation by Beckwith et al. (FPT 2021). We also discuss the required modifications for migrating existing CRYSTALS-Dilithium implementations to match FIPS 204. Through optimized instruction scheduling in the ML-DSA rejection loop, which enables the pre-computation of critical variables, the average signing latency is improved by 16−36%. Quinten Norga, Suparna Kundu, Ingrid Verbauwhede |
DATE | 3 |
| 2026 | A Graph-Theoretic Framework for Randomness Optimization in First-Order Masked CircuitsabstractWe present a generic, automatable framework to reduce the demand for fresh randomness in first-order masked circuits while preserving security in the glitch-extended probing model. The method analyzes the flow of randomness through a circuit to establish security rules based on the glitch-extended probing model. These rules are then encoded as an interference graph, transforming the optimization challenge into a graph coloring problem, which is solved efficiently with a DSATUR heuristic. Crucially, the optimization only rewires randomness inputs without altering core logic, ensuring seamless integration into standard EDA flows and applicability to various gadgets like DOM-indep (Domain-Oriented Masking) and HPC (Hardware Private Circuits). On 32-bit adder architectures, the framework substantially reduces randomness requirements by 79–90%; for instance, the Kogge–Stone adder’s requirement of 259 unique random inputs is reduced to 27. All optimized designs were evaluated using PROLEAD, with the leakage results indicating compliance with first-order glitch-extended probing security. S. V. Dilip Kumar, Benedikt Gierlichs, Ingrid Verbauwhede |
DATE | 3 |
| 2026 | Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory AliasingabstractConfidential computing, powered by trusted execution environments (TEEs) like Intel SGX/TDX and AMD SEV-SNP, is now widely available from major cloud providers. At the core of these technologies is hardware-level memory encryption to protect against privileged attackers and physical threats such as bus snooping and cold boot attacks. Recent extensions add access-control checks to defend against software-based ciphertext manipulation and aliasing attacks. In this work, we challenge the protection modern memory encryption technologies offer against physical adversaries by building a low-cost ($<$\dollarcost) DDR4 interposer that dynamically tampers with address lines to bypass aliasing checks in current TEEs. We demonstrate how the runtime nature of our interposer bypasses boot-time firmware mitigations introduced by AMD and Intel in response to software-based memory aliasing attacks. Using our interposer, we present the first attack on Scalable SGX's single-key domain, achieving arbitrary plaintext read/write access and extracting SGX's platform provisioning key, thereby dismantling trust in remote attestation. We further re-enable a full attestation breach on up-to-date AMD SEV-SNP platforms, bypassing recent firmware defenses against static aliases. Our results challenge core assumptions about encrypted memory security and highlight critical shortcomings in the performance-security trade-offs of current confidential computing systems. Costing orders of magnitude less than commercial DRAM interposers, our device underscores the need for stronger protections against low-cost physical attacks in scalable TEE designs. Jesse De Meulemeester, David F. Oswald, Ingrid Verbauwhede, Jo Van Bulck |
SP | 3 |
| 2026 | A Fully Integrated Quantum Random Number Generator for Cryptographic ApplicationabstractA compact quantum random number generator (QRNG) based on an array of Single Photon Avalanche Diodes (SPADs) is presented here. The main feature of the proposed chip is the capability to generate random numbers without the use of any external source of light. In this view, SPADs are configured to be used in two ways: 1) as a controlled light emitter and 2) as a detector. An embedded logic is able to distinguish dark events against detected emitted photon in such a way to make the system sensitive only to the presence of light. The extracted random bits show a uniform distribution and the QRNG, thanks to an integrated conditioning block maximizing the output entropy, eventually passes all AIS31 tests. The average raw bit rate in the default device configuration (one enabled column) is measured equal to 800 kbps. Nicola Massari, Luca Parmesan, Alessandro Tontini, Sonia Mazzucchi, Milos Grujic, Ingrid Verbauwhede, Andreas Brenneis, Dayo Oshinubi, Ingo Herrmann, Thomas Strohm |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | mUOV: Masking the Unbalanced Oil and Vinegar Digital Signature Scheme at First- and Higher-OrderabstractIn the recent search for additional post-quantum designs, multivariate quadratic equations (MQE) based designs have been receiving attention due to their small signature sizes. Unbalanced Oil and Vinegar (UOV) is an MQE-based digital signature (DS) scheme proposed over two decades ago. Although the mathematical security of UOV has been thoroughly analyzed, several practical side-channel attacks (SCA) have been shown on UOV based DS schemes. In this work, we perform a thorough analysis to identify the variables in UOV based DS schemes that can be exploited with passive SCA, specifically differential power attacks (DPA). Secondly, we introduce masking as a countermeasure to protect the sensitive components of UOV based schemes. We propose efficient masked gadgets for all the critical operations, including the masked dot-product and matrix-vector multiplication. We show that our gadgets are secure in the t-probing model through formal proofs, mechanically verified using the maskVerif tool. We implemented and demonstrated the practical feasibility of our arbitrary-order masking algorithms for UOV-Ip and UOV-III. We show that the masked signature generation of UOV-Ip performs up to 62% better than ML-DSA-44 and 99% better than Falcon-512. In addition, the security of our implementation is practically validated using the test vector leakage assessment (TVLA) methodology. Suparna Kundu, Quinten Norga, Angshuman Karmakar, Uttam Kumar Ojha, Anindya Ganguly, Ingrid Verbauwhede |
CCS | 6 |
| 2025 | Hardware Security: state of the art: KeynoteabstractHardware security is the root of trust in all modern ICT systems!Depending which community you address, it has a different meaning.It covers efficient, secure implementations of new generations of cryptography such as light-weight crypto, post-quantum crypto as well as advanced schemes such as zero-knowledge proofs, fully homomorphic encryption, and computing on encrypted data in general.On top, implementations also must resist a wide variety of side-channel, fault, and micro-architectural attacks.Post-quantum algorithms promise to resist the attacks developed for quantum computers.Yet, their implementations also must resist attacks on classic platforms.Besides crypto, secure systems rely on many more security modules, requiring analog and digital circuit techniques to design quality true random number generators, physically unclonable functions, secure key storage, and many more.A recent report on "Revitalizing the U.S. Semiconductor Ecosystem" (from Executive Office of the President, President's Council of Advisors on Science and Technology, September 2022) describes a set of recommendations on semiconductors and system security.In this presentation, we will demonstrate how our research addresses these recommendations, and we will illustrate this with recent results.This work was supported by Horizon 2020 ERC Advanced Ingrid Verbauwhede |
CF | 1 |
| 2025 | Masking Gaussian Elimination at Arbitrary Order with Application to Multivariate-and Code-Based PQC
Quinten Norga, Suparna Kundu, Uttam Kumar Ojha, Anindya Ganguly, Angshuman Karmakar, Ingrid Verbauwhede |
CT-RSA | 6 |
| 2025 | BadRAM: Practical Memory Aliasing Attacks on Trusted Execution EnvironmentsabstractThe growing adoption of cloud computing raises pressing concerns about trust and data privacy. Trusted Execution Environments (TEEs) have been proposed as promising solutions that implement strong access control and transparent memory encryption within the CPU. While initial TEEs, like Intel SGX, were constrained to small isolated memory regions, the trend is now to protect full virtual machines, e.g., with AMD SEV-SNP, Intel TDX, and Arm CCA. In this paper, we challenge the trust assumptions underlying scaled-up memory encryption and show that an attacker with brief physical access to the embedded SPD chip can cause aliasing in the physical address space, circumventing CPU access control mechanisms. We devise a practical, low-cost setup to create aliases in DDR4 and DDR5 memory modules, breaking the newly introduced integrity guarantees of AMD SEV-SNP. This includes the ability to manipulate memory mappings and corrupt or replay ciphertext, culminating in a devastating end-to-end attack that compromises SEV-SNP's attestation feature. Furthermore, we investigate the issue for other TEEs, demonstrating fine-grained, noiseless write-pattern leakage for classic Intel SGX, while finding that Scalable SGX and TDX employ dedicated alias detection, preventing our attacks at present. In conclusion, our findings dismantle security guarantees in the SEV-SNP ecosystem, necessitating AMD firmware patches, and nuance DRAM trust assumptions for scalable TEE designs. Jesse De Meulemeester, Luca Wilke, David F. Oswald, Thomas Eisenbarth 0001, Ingrid Verbauwhede, Jo Van Bulck |
SP | 5 |
| 2025 | Leuvenshtein: Efficient FHE-based Edit Distance Computation with Single Bootstrap per Cell
Wouter Legiest, Jan-Pieter D'Anvers, Bojan Spasic, Nam-Luc Tran, Ingrid Verbauwhede |
USENIX Security Symposium | 5 |
| 2025 | Scabbard: An Exploratory Study on Hardware Aware Design Choices of Learning with Rounding-based Key Encapsulation MechanismsabstractRecently, the construction of cryptographic schemes based on hard lattice problems has gained immense popularity. Apart from being quantum resistant, lattice-based cryptography allows a wide range of variations in the underlying hard problem. As cryptographic schemes can work in different environments under different operational constraints such as memory footprint, silicon area, efficiency, power requirement, and so on, such variations in the underlying hard problem are very useful for designers to construct different cryptographic schemes. In this work, we explore various design choices of lattice-based cryptography and their impact on performance in the real world. In particular, we propose a suite of key-encapsulation mechanisms based on the learning with rounding problem with a focus on improving different performance aspects of lattice-based cryptography. Our suite consists of three schemes. Our first scheme is Florete, which is designed for efficiency. The second scheme is Espada, which is aimed at improving parallelization, flexibility, and memory footprint. The last scheme is Sable, which can be considered an improved version in terms of key sizes and parameters of the Saber key-encapsulation mechanism, one of the finalists in the National Institute of Standards and Technology’s post-quantum standardization procedure. In this work, we have described our design rationale behind each scheme. Furthermore, to demonstrate the justification of our design decisions, we have provided software and hardware implementations. Our results show Florete is faster than most state-of-the-art KEMs on software platforms. For example, the key-generation algorithm of high-security version Florete outperforms the National Institute of Standards and Technology’s standard Kyber by 47%, the Federal Office for Information Security’s standard Frodo by 99%, and Saber by 57% on the ARM Cortex-M4 platform. Similarly, in hardware, Florete outperforms Frodo and NTRU Prime for all KEM operations. The scheme Espada requires less memory and area than the implementation of most state-of-the-art schemes. For example, the encapsulation algorithm of high-security version Espada uses 30% less stack memory than Kyber, 57% less stack memory than Frodo, and 67% less stack memory than Saber on the ARM Cortex-M4 platform. The implementations of Sable maintain a tradeoff between Florete and Espada regarding software performance and memory requirements. Sable outperforms Saber at least by 6% and Frodo by 99%. Through an efficient polynomial multiplier design, which exploits the small secret size, Sable outperforms most state-of-the-art KEMs, including Saber, Frodo, and NTRU Prime. The implementations of Sable that use number theoretic transform-based polynomial multiplication (SableNTT) surpass all the state-of-the-art schemes in performance, which are optimized for speed on the Cortext M4 platform. The performance benefit of SableNTT against Kyber lies in between 7-29%, 2-13% for Saber, and around 99% for Frodo. Suparna Kundu, Quinten Norga, Angshuman Karmakar, Shreya Gangopadhyay, Jose Maria Bermudo Mera, Ingrid Verbauwhede |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | Low-Cost First-Order Secure Boolean Masking in Glitchy HardwareabstractWe describe how to securely implement the masked logical AND of two bits in hardware in the presence of glitches without the need for fresh randomness, and we provide guidelines for the composition of circuits. As a case study, we design, implement, and evaluate masked DES cores. We focus on first-order secure Boolean masking and do not aim for provable security. Our goal is a practically relevant trade-off between area, latency, randomness cost, and security. We provide two low-cost solutions. Our first solution focuses on strong security while simultaneously aiming for low implementation costs. The resulting DES engine shows no evidence of first-order leakage in a non-specific leakage assessment with 50M traces. Our second solution follows the opposite approach: we focus on lowering implementation costs, latency to be specific, while not sacrificing much on security. Our low-latency DES engine exhibits signs of first-order leakage only after approximately 15M traces. S. V. Dilip Kumar, Josep Balasch, Benedikt Gierlichs, Ingrid Verbauwhede |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | A Practical Key-Recovery Attack on LWE-Based Key-Encapsulation Mechanism Schemes Using Rowhammer
Puja Mondal, Suparna Kundu, Sarani Bhattacharya, Angshuman Karmakar, Ingrid Verbauwhede |
ACNS (3) | 5 |
| 2024 | Hardware Acceleration of the Prime-Factor and Rader NTT for BGV Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) enables computation on encrypted data, holding immense potential for enhancing data privacy and security in various applications. Presently, FHE adoption is hindered by slow computation times, caused by data being encrypted into large polynomials. Optimized FHE libraries and hardware acceleration are emerging to tackle this performance bottleneck. Often, these libraries implement the Number Theoretic Transform (NTT) algorithm for efficient polynomial multiplication. Existing implementations mostly focus on the case where the polynomials are defined over a power-of-two cyclotomic ring, allowing to make use of the simpler Cooley-Tukey NTT. However, generalized cyclotomics have several benefits in the BGV FHE scheme, including more SIMD plaintext slots and a simpler bootstrapping algorithm.We present a hardware architecture for the NTT targeting generalized cyclotomics within the context of the BGV FHE scheme. We explore different non-power-of-two NTT algorithms, including the Prime-Factor, Rader, and Bluestein NTTs. Our most efficient architecture targets the 21845-th cyclotomic polynomial — a practical parameter for BGV — with ideal properties for use with a combination of the Prime-Factor and Rader algorithms. The design achieves high throughput with optimized resource utilization, by leveraging parallel processing, pipelining, and reusing processing elements. Compared to Wu et al.’s VLSI architecture of the Bluestein NTT, our approach showcases 2× to 5× improved throughput and area efficiency. Simulation and implementation results on an AMD Alveo U250 FPGA demonstrate the feasibility of the proposed hardware design for FHE. David Du Pont, Jonas Bertels, Furkan Turan, Michiel Van Beirendonck, Ingrid Verbauwhede |
ARITH | 5 |
| 2024 | A Better Kyber Butterfly for FPGAsabstractKyber was selected by NIST as a Post-Quantum Cryptography Key Encapsulation Mechanism standard. This means that the industry now needs to transition and adopt these new standards. One of the most demanding operations in Kyber is the modular arithmetic, making it a suitable target for optimization. This work offers a novel modular reduction design with the lowest area on Xilinx FPGA platforms. This novel design, through K-reduction and LUT-based reduction, utilizes 49 LUTs and 1 DSP as opposed to Xing and Li’s [XL21] 2021 CHES design requiring 90 LUTs and 1 DSP for one modular multiplication. Our design is the smallest modular multiplier reported as of today. Jonas Bertels, Quinten Norga, Ingrid Verbauwhede |
FPL | 3 |
| 2024 | Reducing Reservoir Dimensionality with Phase Space Construction for Simplified Hardware Implementation
Yuanyang Guo, Robin Degraeve, Philippe Roussel, Ben Kaczer, Erik Bury, Ingrid Verbauwhede |
ICANN (10) | 6 |
| 2024 | Characterization of Oscillator Phase Noise Arising From Multiple Sources for ASIC True Random Number GenerationabstractThis paper presents an analytical study together with an oscillator phase measurement technique to assess the magnitude of the five most prevalent noise types found in free-running oscillators, intended for use in true random number generation. The noise types under study range from white thermal- to flicker- and random walk noise, acting either on the oscillator phase or frequency. A time domain study establishes an analytical connection between the oscillator excess phase variance and the accumulation time length. A variance model that characterizes the phase measurement method is developed, taking into account that one noise type will dominate over all other handled noise types within a specific measurement interval. Additionally, this model considers the effects of both test setup induced variance and quantization noise. The measurement method utilizes a differential approach based on Delay Chains (DCs) and is implemented using a 65nm Complementary Metal-Oxide-Semiconductor (CMOS) technology. A time resolution less than 100 ps could be achieved, effectively lowering the quantization noise floor and allowing to examine jitter accumulation at a time scale down to 30 ns. Measurement results show that typical CMOS ring oscillators are predominantly affected by flicker noise, creating long-term dependencies between the generated period jitter. This observation might in turn invalidate the popular assumption of mutual independence between successive oscillator periods, made in many true random number generator stochastic models. A noise corner is observed at the lower end of the measured accumulation time interval, below which thermal noise becomes dominant over flicker noise, and consecutive period jitter can be considered mutually independent again. Adriaan Peetermans, Ingrid Verbauwhede |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Side-channel Analysis of Lattice-based Post-quantum Cryptography: Exploiting Polynomial MultiplicationabstractPolynomial multiplication algorithms such as Toom-Cook and the Number Theoretic Transform are fundamental building blocks for lattice-based post-quantum cryptography. In this work we present correlation power-analysis-based side-channel analysis methodologies targeting every polynomial multiplication strategy for all lattice-based post-quantum key encapsulation mechanisms in the final round of the NIST post-quantum standardization procedure. We perform practical experiments on real side-channel measurements, demonstrating that our method allows to extract the secret key from all lattice-based post-quantum key encapsulation mechanisms. Our analysis shows that the used polynomial multiplication strategy can significantly impact the time complexity of the attack. Catinca Mujdei, Lennert Wouters, Angshuman Karmakar, Arthur Beckers, Jose Maria Bermudo Mera, Ingrid Verbauwhede |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2024 | Optimizing Linear Correctors: A Tight Output Min-Entropy Bound and Selection TechniqueabstractPost-processing of the raw bits produced by a true random number generator (TRNG) is always necessary when the entropy per bit is insufficient for security applications. In this paper, we derive a tight bound on the output min-entropy of the algorithmic post-processing module based on linear codes, known as linear correctors. Our bound is based on the codes’ weight distributions, and we prove that it holds even for the real-world noise sources that produce independent but not identically distributed bits. Additionally, we present a method for identifying the optimal linear corrector for a given input min-entropy rate that maximizes the throughput of the post-processed bits while simultaneously achieving the needed security level. Our findings show that for an output min-entropy rate of 0.999, the extraction efficiency of the linear correctors with the new bound can be up to$\mathbf {130.56\, \%}$higher when compared to the old bound, with an average improvement of$\mathbf {41.2\, \%}$over the entire input min-entropy range. On the other hand, the required min-entropy of the raw bits for the individual correctors can be reduced by up to$\mathbf {61.62\, \%}$. Milos Grujic, Ingrid Verbauwhede |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | ShowTime: Amplifying Arbitrary CPU Timing Side ChannelsabstractMicroarchitectural attacks typically rely on precise timing sources to uncover short-lived secret-dependent activity in the processor. In response, many browsers and even CPU vendors restrict access to fine-grained timers. While some attacks are still possible, several state-of-the-art microarchitectural attack vectors are actively hindered or even eliminated by these restrictions. Antoon Purnal, Marton Bognar, Frank Piessens, Ingrid Verbauwhede |
AsiaCCS | 4 |
| 2023 | An In-Depth Security Evaluation of the Nintendo DSi Gaming Console
pcy Sluys, Lennert Wouters, Benedikt Gierlichs, Ingrid Verbauwhede |
CARDIS | 4 |
| 2023 | FPT: A Fixed-Point Accelerator for Torus Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) is a technique that allows computation on encrypted data. It has the potential to drastically change privacy considerations in the cloud, but high computational and memory overheads are preventing its broad adoption. TFHE is a promising Torus-based FHE scheme that heavily relies on bootstrapping, the noise-removal tool invoked after each encrypted logical/arithmetical operation. Michiel Van Beirendonck, Jan-Pieter D'Anvers, Furkan Turan, Ingrid Verbauwhede |
CCS | 4 |
| 2023 | On the Unpredictability of SPICE Simulations for Side-Channel Leakage Verification of Masked Cryptographic CircuitsabstractCircuits for cryptography are vulnerable to side-channel (SC) attacks. Masking is a countermeasure which splits secrets into random shares. It is provable secure under the assumption that physical leakage of each share is independent of each other. For a secure implementation of masked circuits, this independency assumption must be satisfied after layout. A transistor-level simulator such as SPICE produces analog waveforms that are sufficiently trustworthy to verify timing accuracy. Due to this accuracy, SPICE is expected to be useful for SC leakage verification after layout. However, we demonstrate that the statistical variation of the power noise amplitude in SPICE simulation is not always correct and varies a lot for SC evaluation. We believe it results from the internal time-step creation optimized for efficiency. It causes false-positives in the verification of security order. A small nonlinear function with a domain-oriented masking scheme is used to demonstrate these SPICE-simulation anomalies. Kazuki Monta, Makoto Nagata, Josep Balasch, Ingrid Verbauwhede |
DAC | 4 |
| 2023 | Low-Cost First-Order Secure Boolean Masking in Glitchy HardwareabstractWe describe how to securely implement the logical AND of two bits in hardware in the presence of glitches without the need for fresh randomness. As a case study, we design, implement and evaluate a DES core using our AND gate. Our goal is an overall practically relevant tradeoff between area, latency, randomness cost and security. We focus on first-order secure Boolean masking and we do not aim for provable security. The resulting DES engine shows no evidence of first-order leakage in a non-specific leakage assessment with 50M traces. S. V. Dilip Kumar, Josep Balasch, Benedikt Gierlichs, Ingrid Verbauwhede |
DATE | 4 |
| 2023 | Hardware Acceleration of FHEWabstractThe magic of Fully Homomorphic Encryption (FHE) is that it allows operations on encrypted data without decryption. Unfortunately, the slow computation time limits their adoption. The slow computation time results from the vast memory requirements (64Kbits per ciphertext), a bootstrapping key of 1.3 GB, and sizeable computational overhead (10240 NTTs, each NTT requiring 5120 32-bit multiplications). We accelerate the FHEW bootstrapping in hardware on a high-end U280 FPGA.To reduce the computational complexity, we propose a fast hardware NTT architecture modified from [5] with support for negatively wrapped convolution. The IP module includes large I/O ports to the NTT accelerator and an index bit-reversal block. The total architecture requires less than 225000 LUTs and 1280 DSPs.Assuming that a fast interface to the FHEW bootstrapping key is available, the execution speed of FHEW bootstrapping can increase by at least 7.5 times. Jonas Bertels, Michiel Van Beirendonck, Furkan Turan, Ingrid Verbauwhede |
DDECS | 4 |
| 2023 | Neural Network Quantisation for Faster Homomorphic EncryptionabstractHomomorphic encryption (HE) enables calculating on encrypted data, which makes it possible to perform privacy-preserving neural network inference. One disadvantage of this technique is that it is several orders of magnitudes slower than calculation on unencrypted data. Neural networks are commonly trained using floating-point, while most homomorphic encryption libraries calculate on integers, thus requiring a quantisation of the neural network. A straightforward approach would be to quantise to large integer sizes (e.g. 32 bit) to avoid large quantisation errors. In this work, we reduce the integer sizes of the networks, using quantisation-aware training, to allow more efficient computations. For the targeted MNIST architecture proposed by Badawi et al. [1], we reduce the integer sizes by 33% without significant loss of accuracy, while for the CIFAR architecture, we can reduce the integer sizes by 43%. Implementing the resulting networks under the BFV homomorphic encryption scheme using SEAL, we could reduce the execution time of an MNIST neural network by 80% and by 40% for a CIFAR neural network. Wouter Legiest, Furkan Turan, Michiel Van Beirendonck, Jan-Pieter D'Anvers, Ingrid Verbauwhede |
IOLTS | 5 |
| 2023 | SpectrEM: Exploiting Electromagnetic Emanations During Transient Execution
Jesse De Meulemeester, Antoon Purnal, Lennert Wouters, Arthur Beckers, Ingrid Verbauwhede |
USENIX Security Symposium | 5 |
| 2023 | Revisiting Higher-Order Masked Comparison for Lattice-Based Cryptography: Algorithms and Bit-Sliced ImplementationsabstractMasked comparison is one of the most expensive operations in side-channel secure implementations of lattice-based post-quantum cryptography, especially for higher masking orders. First, we introduce two new masked comparison algorithms, which improve the arithmetic comparison of D’Anvers et al. (2021) and the hybrid comparison method of Coron et al. (2021) respectively. We then look into implementation-specific optimizations, and show that small specific adaptations can have a significant impact on the overall performance. Finally, we implement various state-of-the-art comparison algorithms and benchmark them on the same platform (ARM-Cortex M4) to allow a fair comparison between them. We improve on the arithmetic comparison of D’Anvers et al. with a factor$\approx 20\%$by using Galois Field multiplications and the hybrid comparison of Coron et al. with a factor$\approx 25\%$by streamlining the design. Our implementation-specific improvements allow a speedup of a straightforward comparison implementation of$\approx 33\%$. We discuss the differences between the various algorithms and provide the implementations and a testing framework to ease future research. Jan-Pieter D'Anvers, Michiel Van Beirendonck, Ingrid Verbauwhede |
IEEE Trans. Computers | 3 |
| 2022 | Mining CryptoNight-Haven on the Varium C1100 Blockchain Accelerator CardabstractCryptocurrency mining is an energy-intensive process that presents a prime candidate for hardware acceleration. This work-in-progress presents the first coprocessor design for the ASIC-resistant CryptoNight-Haven Proof of Work (PoW) algorithm. We construct our hardware accelerator as a Xilinx Run Time (XRT) RTL kernel targeting the Xilinx Varium C1100 Blockchain Accelerator Card. The design employs deeply pipelined computation and High Bandwidth Memory (HBM) for the underlying scratchpad data. We aim to compare our accelerator to existing CPU and GPU miners to show increased throughput and energy efficiency of its hash computations. Lucas Bex, Furkan Turan, Michiel Van Beirendonck, Ingrid Verbauwhede |
FPL | 4 |
| 2022 | Hardware Security: Physical Design versus Side-Channel and Fault AttacksabstractWhat is "hardware" security? How can we improve trustworthiness in hardware circuits? Is there a design method for secure hardware design? To answer these questions, different communities have different expectations of trusted (expecting trustworthy) hardware components upon which they start to build a secure system. At the same time, electronics shrink: sensor nodes, IOT devices, smart electronics are becoming more and more available. In the past, adding security was only a concern for locked server rooms or now cloud servers. However, these days, our portable devices contain highly private and secure information. Adding security and cryptography to these often very resource constraint devices is a challenge. Moreover, they can be subject to physical attacks, including side-channel and fault attacks [1][2]. This presentation aims at bringing some order in the chaos of expectations by introducing the importance of a design methodology for secure design [3][5]. We will illustrate the capabilities of current side EM and laser fault passive and active attacks. In this context, we will also reflect on the role of physical design, place and route [4][6]. Ingrid Verbauwhede |
ISPD | 1 |
| 2022 | Double Trouble: Combined Heterogeneous Attacks on Non-Inclusive Cache Hierarchies
Antoon Purnal, Furkan Turan, Ingrid Verbauwhede |
USENIX Security Symposium | 3 |
| 2022 | TROT: A Three-Edge Ring Oscillator Based True Random Number Generator With Time-to-Digital ConversionabstractThis paper introduces a new true random number generator (TRNG) based on a three-edge ring oscillator. Our design uses a new technique with a time-to-digital converter to effectively acquire jitter accumulated independently by each edge. As a part of the security evaluation, we present the stochastic model of the TRNG’s digital noise source and estimate a lower bound of the min-entropy per random bit. Starting from the obtained entropy bound, we propose a procedure for selecting and implementing an area-efficient and throughput-optimal post-processing function based on the best known linear codes that will increase the output min-entropy rate to more than 0.999. The proposed TRNG exquisitely balances low design effort and resource consumption with high throughput and a high min-entropy rate, making it more suitable for randomness-demanding and resource-constrained platforms than the state-of-the-art. The complete implementation of the TRNG digital noise source and the post-processing occupies 33 slices and achieves a throughput of 12.5 Mbps on Xilinx Zynq-7000 FPGAs. The min-entropy of the generated random bits is assessed by NIST SP 800-90B entropy estimators, and the tested sequences pass the AIS-31 test suit. Milos Grujic, Ingrid Verbauwhede |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Prime+Scope: Overcoming the Observer Effect for High-Precision Cache Contention AttacksabstractModern processors expose software to information leakage through shared microarchitectural state. One of the most severe leakage channels is cache contention, exploited by attacks referred to as PRIME+PROBE, which can infer fine-grained memory access patterns while placing only limited assumptions on attacker capabilities. Antoon Purnal, Furkan Turan, Ingrid Verbauwhede |
CCS | 3 |
| 2021 | Exploring Micro-architectural Side-Channel Leakages through Statistical TestingabstractMicro-architectural side-channel leakage received a lot of attention due to their high impact on software security on complex out-of-order processors. These are extremely specialised threat models and can be only realised in practise with high precision measurement code, triggering micro-architectural behavior that leaks information. In this paper, we present a tool to support the inexperienced user to verify his code for side-channel leakage. We combine two very useful tools- statistical testing and hardware performance monitors to bridge this gap between the understanding of the general purpose users and the most precise speculative execution attacks. We first show that these event counters are more powerful than observing timing variabilities on an executable. We extend Dudect, where the raw hardware events are collected over the target executable, and leakage detection tests are incorporated on the statistics of observed events following the principles of non-specific t-tests. Finally, we show the applicability of our tool on the most popular speculative micro-architectural and data-sampling attack models. Sarani Bhattacharya, Ingrid Verbauwhede |
DATE | 2 |
| 2021 | Systematic Analysis of Randomization-based Protected Cache ArchitecturesabstractRecent secure cache designs aim to mitigate side-channel attacks by randomizing the mapping from memory addresses to cache sets. As vendors investigate deployment of these caches, it is crucial to understand their actual security.In this paper, we consolidate existing randomization-based secure caches into a generic cache model. We then comprehensively analyze the security of existing designs, including CEASER-S and SCATTERCACHE, by mapping them to instances of this model. We tailor cache attacks for randomized caches using a novel PRIME+PRUNE+PROBE technique, and optimize it using burst accesses, bootstrapping, and multi-step profiling. PRIME+ PRUNE+PROBE constructs probabilistic but reliable eviction sets, enabling attacks previously assumed to be computationally infeasible. We also simulate an end-to-end attack, leaking secrets from a vulnerable AES implementation. Finally, a case study of CEASER-S reveals that cryptographic weaknesses in the randomization algorithm can lead to a complete security subversion.Our systematic analysis yields more realistic and comparable security levels for randomized caches. As we quantify how design parameters influence the security level, our work leads to important conclusions for future work on secure cache designs. Antoon Purnal, Lukas Giner, Daniel Gruss, Ingrid Verbauwhede |
SP | 4 |
| 2021 | A Side-Channel-Resistant Implementation of SABERabstractThe candidates for the NIST Post-Quantum Cryptography standardization have undergone extensive studies on efficiency and theoretical security, but research on their side-channel security is largely lacking. This remains a considerable obstacle for their real-world deployment, where side-channel security can be a critical requirement. This work describes a side-channel-resistant instance of Saber, one of the lattice-based candidates, using masking as a countermeasure. Saber proves to be very efficient to masking due to two specific design choices: power-of-two moduli and limited noise sampling of learning with rounding. A major challenge in masking lattice-based cryptosystems is the integration of bit-wise operations with arithmetic masking, requiring algorithms to securely convert between masked representations. The described design includes a novel primitive for masked logical shifting on arithmetic shares and adapts an existing masked binomial sampler for Saber. An implementation is provided for an ARM Cortex-M4 microcontroller, and its side-channel resistance is experimentally demonstrated. The masked implementation features a 2.5x overhead factor, significantly lower than the 5.7x previously reported for a masked variant of NewHope. Masked key decapsulation requires less than 3,000,000 cycles on the Cortex-M4 and consumes less than 12kB of dynamic memory, making it suitable for deployment in embedded platforms. Michiel Van Beirendonck, Jan-Pieter D'Anvers, Angshuman Karmakar, Josep Balasch, Ingrid Verbauwhede |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2021 | Design and Analysis of Configurable Ring Oscillators for True Random Number Generation Based on Coherent SamplingabstractTrue Random Number Generators (TRNGs) are indispensable in modern cryptosystems. Unfortunately, to guarantee high entropy of the generated numbers, many TRNG designs require a complex implementation procedure, often involving manual placement and routing. In this work, we introduce, analyse, and compare three dynamic calibration mechanisms for the COherent Sampling ring Oscillator based TRNG: GateVar , WireVar , and LUTVar , enabling easy integration of the entropy source into complex systems. The TRNG setup procedure automatically selects a configuration that guarantees the security requirements. In the experiments, we show that two out of the three proposed mechanisms are capable of assuring correct TRNG operation even when an automatic placement is carried out and when the design is ported to another Field-Programmable Gate Array (FPGA) family. We generated random bits on both a Xilinx Spartan 7 and a Microsemi SmartFusion2 implementation that, without post processing, passed the AIS-31 statistical tests at a throughput of 4.65 Mbit/s and 1.47 Mbit/s, respectively. Adriaan Peetermans, Vladimir Rozic, Ingrid Verbauwhede |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2020 | Compact domain-specific co-processor for accelerating module lattice-based KEMabstractWe present a domain-specific co-processor to speed up Saber, a post-quantum key encapsulation mechanism competing on the NIST Post-Quantum Cryptography standardization process. Contrary to most lattice-based schemes, Saber doesn’t use NTT-based polynomial multiplication. We follow a hardware-software co-design approach: the execution is performed on an ARM core and only the most computationally expensive operation, i.e., the polynomial multiplication, is offloaded to the co-processor to obtain a compact design. We exploit the idea of distributed computing at micro-architectural level together with novel algorithmic optimizations to achieve approximately a 6 times speedup with respect to optimized software at a small area cost, which we demonstrate on a Zynq-7000 ARM/FPGA SoC. Jose Maria Bermudo Mera, Furkan Turan, Angshuman Karmakar, Sujoy Sinha Roy, Ingrid Verbauwhede |
DAC | 5 |
| 2020 | Sweeping for Leakage in Masked Circuit LayoutsabstractMasking schemes are the most popular countermeasure against side-channel analysis. They theoretically decorrelate information leaked through inherent physical channels from the key-dependent intermediate values that occur during computation. Their provable security is devised under models that abstract complex physical phenomena of the underlying hardware. In this work, we investigate the impact of the physical layout to the side-channel security of masking schemes. For this we propose a model for co-simulation of the analog power distribution network with the digital logic core. Our study considers the drive of the power supply buffers, as well as parasitic resistors, inductors and capacitors. We quantify our findings using Test Vector Leakage Assessment by relative comparison to the parasitic-free model. Thus we provide a deeper insight into the potential layout sources of leakage and their magnitude. Danilo Sijacic, Josep Balasch, Ingrid Verbauwhede |
DATE | 3 |
| 2020 | Attacking Hardware Random Number Generators in a Multi-Tenant ScenarioabstractTrue random number generators are important building blocks for cryptographic systems and can be the target of adversaries that want to break cryptographic protocols by reducing the unpredictability of the used random numbers. This paper examines the viability of three different types of potential attacks on these generators when they are implemented on field programmable gate arrays, namely the voltage manipulation attack, the ring-oscillator locking attack and the replica observation attack. The proposed attacks only make use of the available programmable logic of the device and as such do not require physical access to it. They can technically be mounted remotely in a multi-tenant scenario by adversaries that only have bitstream write access to a part of the programmable logic. The attacks try to exploit interactions that can exist between an attack circuit and the targeted circuit because they reside on the same chip. The paper presents two case studies: an elementary ring oscillator design and a transition effect ring oscillator design. For the first case study, all three scenarios were tested and for the second case study, only the voltage manipulation attack scenario is examined. Our results show that this voltage manipulation attack is the most effective of the three proposed attacks. Yrjo Koyen, Adriaan Peetermans, Vladimir Rozic, Ingrid Verbauwhede |
FDTC | 4 |
| 2020 | HEAWS: An Accelerator for Homomorphic Encryption on the Amazon AWS FPGAabstractHomomorphic Encryption makes privacy preserving computing possible in a third party owned cloud by enabling computation on the encrypted data of users. However, software implementations of homomorphic encryption are very slow on general purpose processors. With the emergence of `FPGAs as a service', hardware-acceleration of computationally heavy workloads in the cloud are getting popular. In this article we propose HEAWS, a domain-specific coprocessor architecture for accelerating homomorphic function evaluation on the encrypted data using high-performance FPGAs available in the Amazon AWS cloud. To the best of our knowledge, we are the first to report hardware acceleration of homomorphic encryption using Amazon AWS FPGAs. Utilizing the massive size of the AWS FPGAs, we design a high-performance and parallel coprocessor architecture for the FV homomorphic encryption scheme which has become popular for computing exact arithmetic on the encrypted data. We design parallel building blocks and apply pipeline processing at different levels of the implementation hierarchy, and on top of such optimizations we instantiate multiple parallel coprocessors in the FPGA to execute several homomorphic computations simultaneously. While the absolute computation time can be reduced by deploying more computational resources, efficiency of the HW/SW communication interface plays an important role in homomorphic encryption as it is computation as well as data intensive. Our implementation utilizes state of the art 512-bit XDMA feature of high bandwidth communication available in the AWS Shell to reduce the overhead of HW/SW data transfer. Moreover, we explore the design-space to identify optimal off-chip data transfer strategy for feeding the parallel coprocessors in a time-shared manner. As a result of these optimizations, our AWS-based accelerator can perform 613 homomorphic multiplications per second for a parameter set that enables homomorphic computations of depth 4. Finally, we benchmark an artificial neural network for privacy-preserving forecasting of energy consumption in a Smart Grid application and observe five times speed up. Furkan Turan, Sujoy Sinha Roy, Ingrid Verbauwhede |
IEEE Trans. Computers | 3 |
| 2019 | Design Considerations for EM Pulse Fault Injection
Arthur Beckers, Masahiro Kinugawa, Yuichi Hayashi, Daisuke Fujimoto, Josep Balasch, Benedikt Gierlichs, Ingrid Verbauwhede |
CARDIS | 7 |
| 2019 | Design Principles for True Random Number Generators for Security ApplicationsabstractThe generation of high quality true random numbers is essential in security applications. For secure communication, we also require high quality true random number generators (TRNGs) in embedded and IoT devices. This paper provides insights into modern TRNG design principles and their evaluation, based on standard's requirements and design experience. We illustrate our approach with a case study of a recently proposed delay chain based TRNG. Milos Grujic, Vladimir Rozic, David Johnston, John Kelsey, Ingrid Verbauwhede |
DAC | 5 |
| 2019 | Pushing the speed limit of constant-time discrete Gaussian sampling. A case study on the Falcon signature schemeabstractSampling from a discrete Gaussian distribution has applications in lattice-based post-quantum cryptography. Several efficient solutions have been proposed in recent years. However, making a Gaussian sampler secure against timing attacks turned out to be a challenging research problem. In this work, we present a toolchain to instantiate an efficient constant-time discrete Gaussian sampler of arbitrary standard deviation and precision. We observe an interesting property of the mapping from input random bit strings to samples during a Knuth-Yao sampling algorithm and propose an efficient way of minimizing the Boolean expressions for the mapping. Our minimization approach results in up to 37% faster discrete Gaussian sampling compared to the previous work. Finally, we apply our optimized and secure Gaussian sampler in the lattice-based digital signature algorithm Falcon, which is a NIST submission, and provide experimental evidence that the overall performance of the signing algorithm degrades by at most 33% only due to the additional overhead of 'constant-time' sampling, including the 60% overhead of random number generation. Breaking a general belief, our results indirectly show that the use of discrete Gaussian samples in digital signature algorithms would be beneficial. Angshuman Karmakar, Sujoy Sinha Roy, Frederik Vercauteren, Ingrid Verbauwhede |
DAC | 4 |
| 2019 | A Self-Calibrating True Random Number GeneratorabstractTrue Random Number Generators (TRNGs) are essential in all security systems. Unfortunately, large design effort is required to ensure that a TRNG design on a Field-Programmable Gate Array (FPGA) generates a sufficient entropy density at its output. This design effort relates to the fact that for each FPGA family a manual placement and routing procedure has to be executed. On top of this often comes the additional effort of finding a suitable location inside the target FPGA. This searching procedure has to be repeated for every device separately. In this demo, we show the working of a novel entropy source for the Coherent Sampling Ring Oscillator (COSO) based TRNG. This entropy source eliminates the need for any manual intervention during the implementation process. It generates two oscillating signals that can be matched with a precision of a few picoseconds. A controller regulates this entropy source based on some predefined bounds on the period length difference of the two oscillating signals. Adriaan Peetermans, Milos Grujic, Vladimir Rozic, Ingrid Verbauwhede |
FPL | 4 |
| 2019 | A Highly-Portable True Random Number Generator Based on Coherent SamplingabstractTrue Random Number Generators (TRNGs) are indispensable in modern cryptosystems. Unfortunately, in order to guarantee high entropy of the generated numbers, many TRNG designs require a complex implementation procedure, often involving manual placement and routing. In this work, we introduce a dynamic calibration mechanism for the Coherent Sampling Ring Oscillator based TRNG (COSO-TRNG) enabling easy integration of the entropy source into complex systems. The TRNG setup procedure automatically selects a configuration that guarantees the security requirements. In the experiments, we show that the proposed mechanism is capable of assuring correct TRNG operation even when an automatic placement is carried out and when the design is ported to another FPGA family. We generated random bits on both a Xilinx Spartan 6 and a Microsemi SmartFusion2 implementation that, without post processing, passed AIS-31 statistical tests at a throughput of 3.30 Mbit/s and 1.47 Mbit/s respectively. Adriaan Peetermans, Vladimir Rozic, Ingrid Verbauwhede |
FPL | 3 |
| 2019 | FPGA-Based High-Performance Parallel Architecture for Homomorphic Computing on Encrypted DataabstractHomomorphic encryption is a tool that enables computation on encrypted data and thus has applications in privacy-preserving cloud computing. Though conceptually amazing, implementation of homomorphic encryption is very challenging and typically software implementations on general purpose computers are extremely slow. In this paper we present our year long effort to design a domain specific architecture in a heterogeneous Arm+FPGA platform to accelerate homomorphic computing on encrypted data. We design a custom co-processor for the computationally expensive operations of the well-known Fan-Vercauteren (FV) homomorphic encryption scheme on the FPGA, and make the Arm processor a server for executing different homomorphic applications in the cloud, using this FPGA-based co-processor. We use the most recent arithmetic and algorithmic optimization techniques and perform designspace exploration on different levels of the implementation hierarchy. In particular we apply circuit-level and block-level pipeline strategies to boost the clock frequency and increase the throughput respectively. To reduce computation latency, we use parallel processing at all levels. Starting from the highly optimized building blocks, we gradually build our multi-core multi-processor architecture for computing. We implemented and tested our optimized domain specific programmable architecture on a single Xilinx Zynq UltraScale+ MPSoC ZCU102 Evaluation Kit. At 200 MHz FPGA-clock, our implementation achieves over 13x speedup with respect to a highly optimized software implementation of the FV homomorphic encryption scheme on an Intel i5 processor running at 1.8 GHz. Sujoy Sinha Roy, Furkan Turan, Kimmo Järvinen 0001, Frederik Vercauteren, Ingrid Verbauwhede |
HPCA | 5 |
| 2019 | The Impact of Error Dependencies on Ring/Mod-LWE/LWR Based Schemes
Jan-Pieter D'Anvers, Frederik Vercauteren, Ingrid Verbauwhede |
PQCrypto | 3 |
| 2019 | Single-Round Pattern Matching Key Generation Using Physically Unclonable FunctionabstractParal and Devadas introduced a simple key generation scheme with a physically unclonable function (PUF) that requires no error correction, e.g., by using a fuzzy extractor. Their scheme, called a pattern matching key generation (PMKG) scheme, is based on pattern matching between auxiliary data, assigned at the enrollment in advance, and a substring of PUF output, to reconstruct a key. The PMKG scheme repeats a round operation, including the pattern matching, to derive a key with high entropy. Later, to enhance the efficiency and security, a circular PMKG (C-PMKG) scheme was proposed. However, multiple round operations in these schemes make them impractical. In this paper, we propose a single-round circular PMKG (SC-PMKG) scheme. Unlike the previous schemes, our scheme invokes the PUF only once. Hence, there is no fear of information leakage by invoking the PUF with the (partially) same input multiple times in different rounds, and, therefore, the security consideration can be simplified. Moreover, we introduce another hash function to generate a check string which ensures the correctness of the key reconstruction. The string enables us not only to defeat manipulation attacks but also to prove the security theoretically. In addition to its simple construction, the SC-PMKG scheme can use a weak PUF like the SRAM-PUF as a building block if our system is properly implemented so that the PUF is directly inaccessible from the outside, and, therefore, it is suitable for tiny devices in the IoT systems. We discuss its security and show its feasibility by simulations and experiments. Yuichi Komano, Kazuo Ohta, Kazuo Sakiyama, Mitsugu Iwamoto, Ingrid Verbauwhede |
Secur. Commun. Networks | 5 |
| 2019 | Atlas: Application Confidentiality in Compromised Embedded SystemsabstractDue to the requirements of the Internet-of-Things, modern embedded systems have become increasingly complex, running different applications. In order to protect their intellectual property as well as the confidentiality of sensitive data they process, these applications have to be isolated from each other. Traditional memory protection and memory management units provide such isolation, but rely on operating system support for their configuration. However, modern operating systems tend to be vulnerable and cannot guarantee confidentiality when compromised. We present Atlas, a hardware-based security architecture, complementary to traditional memory protection mechanisms, ensuring code and data confidentiality through transparent encryption, even when the system software has been exploited. Atlas relies on its zero-software trusted computing base to protect against system-level attackers and also supports secure shared memory. We implemented Atlas based on the LEON3 softcore processor, including toolchain extensions for developers. Our FPGA-based evaluation shows minimal cycle overhead at the cost of a reduced maximum frequency. Pieter Maene, Johannes Götzfried, Tilo Müller, Ruan de Clercq, Felix C. Freiling, Ingrid Verbauwhede |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2019 | Compact and Flexible FPGA Implementation of Ed25519 and X25519abstractThis article describes a field-programmable gate array (FPGA) cryptographic architecture, which combines the elliptic curve--based Ed25519 digital signature algorithm and the X25519 key establishment scheme in a single module. Cryptographically, these are high-security elliptic curve cryptography algorithms with short key sizes and impressive execution times in software. Our goal is to provide a lightweight FPGA module that enables them on resource-constrained devices, specifically for Internet of Things (IoT) applications. In addition, we aim at extensibility with customisable countermeasures against timing and differential power analysis side-channel attacks and fault-injection attacks. For the former, we offer a choice between time-optimised versus constant-time execution, with or without Z -coordinate randomisation and base-point blinding; and for the latter, we offer enabling or disabling default-case statements in the Finite State Machine (FSM) descriptions. To obtain compactness and at the same time fast execution times, we make maximum use of the Digital Signal Processing (DSP) slices on the FPGA. We designed a single arithmetic unit that is flexible to support operations with two moduli and non-modulus arithmetic. In addition, our design benefits in-place memory management and the local storage of inputs into DSP slices’ pipeline registers and takes advantage of distributed memory. These eliminate a memory access bottleneck. The flexibility is offered by a micro-code supported instruction-set architecture. Our design targets 7-Series Xilinx FPGAs and is prototyped on a Zynq System-on-Chip (SoC). The base design combining Ed25519 and X25519 in a single module, and its implementation requires only around 11.1K Lookup Tables (LUTs), 2.6K registers, and 16 DSP slices. Also, it achieves performance of 1.6ms for a signature generation and 3.6ms for a signature verification for a 1024-bit message with an 82MHz clock. Moreover, the design can be optimised only for X25519, which gives the most compact FPGA implementation compared to previously published X25519 implementations. Furkan Turan, Ingrid Verbauwhede |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | An In-Depth and Black-Box Characterization of the Effects of Laser Pulses on ATmega328P
S. V. Dilip Kumar, Arthur Beckers, Josep Balasch, Benedikt Gierlichs, Ingrid Verbauwhede |
CARDIS | 5 |
| 2018 | Towards inter-vendor compatibility of true random number generators for FPGAsabstractTrue random number generators (TRNGs) are fundamental constituents of secure embedded cryptographic systems. In this paper, we introduce a general methodology for porting TRNG across different FPGA vendor families. In order to demonstrate our methodology, we applied it to the delay-chain based TRNG (DC-TRNG) on Intel Cyclone IV and Cyclone V FPGAs. We examine vendor-agnostic generality of the underlying DC-TRNG principle and propose modifications to address differences in structure of FPGAs. Implementation of the DC-TRNG on Cyclone IV uses 149 LEs (<;0.1% of available resources) and has a throughput of 5Mbps, while on Cyclone V it occupies 230 ALMs (<;1.5% of resources) with an output rate of 12.5 Mbps. The quality of the random bits produced by the DC-TRNG on Intel Cyclone IV and V is further confirmed by using NIST statistical test suite. Milos Grujic, Bohan Yang 0001, Vladimir Rozic, Ingrid Verbauwhede |
DATE | 4 |
| 2018 | Design and testing methodologies for true random number generators towards industry certificationabstractThe objective of this paper is to provide insight on the design, evaluation and testing of modern True Random Number Generators (TRNGs) aimed towards certification. We discuss aspects related to each of these stages by means of two illustrative TRNG designs: PLL-TRNG and DC-TRNG. Topics covered in the paper include: the importance of formal security evaluations based on a stochastic model of the entropy source, the development of suitable and lightweight embedded tests to detect failures, the implementation and testing of TRNGs in dedicated FPGA platforms, and a robustness assessment to environmental and/or physical modifications. Josep Balasch, Florent Bernard, Viktor Fischer, Milos Grujic, Marek Laban, Oto Petura, Vladimir Rozic, Gerard van Battum, Ingrid Verbauwhede, Marnix Wakker, Bohan Yang 0001 |
ETS | 9 |
| 2018 | The Impact of Pulsed Electromagnetic Fault Injection on True Random Number GeneratorsabstractRandom number generation is a key function of today's secure devices. Commonly used for key generation, random number streams are more and more frequently used as the anchor of trust of several countermeasures such as masking. True Random Number Generators (TRNGs) thus become a relevant entry point for attacks that aim at lowering the security of integrated systems. Within this context, this paper investigates the robustness of TRNGs based on Ring Oscillators (focusing on the delay chain TRNG) against pulsed electromagnetic fault injection. Indeed, weaknesses in generating random bits for masking scheme degenerate the Side Channel resistance. Finally by exploiting fault results on delay chain TRNG some general guidelines to harden them are derived. Maxime Madau, Michel Agoyan, Josep Balasch, Milos Grujic, Patrick Haddad, Philippe Maurine, Vladimir Rozic, Dave Singelée, Bohan Yang 0001, Ingrid Verbauwhede |
FDTC | 10 |
| 2018 | A Closer Look at the Delay-Chain based TRNGabstractThis paper presents a refined stochastic model of the delay-chain based true random number generator (DC-TRNG) and its application. DC-TRNG is a true random number generator for FPGAs that utilizes time-to-digital conversion (TDC) to accurately determine the position of the ring-oscillator jittery signal edge. Our stochastic model employs precise time characterization of the carry-chains that are used for TDC in the DC-TRNG. In order to determine lower bounds of the estimated min-entropy, the binary probabilities are calculated by applying the stochastic model. Based on these computed probabilities, we perform optimizations of the DC-TRNG parameters on two different FPGAs - Xilinx Spartan 6 and Intel Cyclone IV, in order to achieve the highest possible throughput of the DC-TRNG. Milos Grujic, Vladimir Rozic, Bohan Yang 0001, Ingrid Verbauwhede |
ISCAS | 4 |
| 2018 | Constant-Time Discrete Gaussian SamplingabstractSampling from a discrete Gaussian distribution is an indispensable part of lattice-based cryptography. Several recent works have shown that the timing leakage from a non-constant-time implementation of the discrete Gaussian sampling algorithm could be exploited to recover the secret. In this paper, we propose a constant-time implementation of the Knuth-Yao random walk algorithm for performing constant-time discrete Gaussian sampling. Since the random walk is dictated by a set of input random bits, we can express the generated sample as a function of the input random bits. Hence, our constant-time implementation expresses the unique mapping of the input random-bits to the output sample-bits as a Boolean expression of the random-bits. We use bit-slicing to generate multiple samples in batches and thus increase the throughput of our constant-time sampling manifold. Our experiments on an Intel i7-Broadwell processor show that our method can be as much as 2.4 times faster than the constant-time implementation of cumulative distribution table based sampling and consumes exponentially less memory than the Knuth-Yao algorithm with shuffling for a similar level of security. Angshuman Karmakar, Sujoy Sinha Roy, Oscar Reparaz, Frederik Vercauteren, Ingrid Verbauwhede |
IEEE Trans. Computers | 5 |
| 2018 | Hardware-Based Trusted Computing Architectures for Isolation and AttestationabstractAttackers target many different types of computer systems in use today, exploiting software vulnerabilities to take over the device and make it act maliciously. Reports of numerous attacks have been published, against the constrained embedded devices of the Internet of Things, mobile devices like smartphones and tablets, high-performance desktop and server environments, as well as complex industrial control systems. Trusted computing architectures give users and remote parties like software vendors guarantees about the behaviour of the software they run, protecting them against software-level attackers. This paper defines the security properties offered by them, and presents detailed descriptions of twelve hardware-based attestation and isolation architectures from academia and industry. We compare all twelve designs with respect to the security properties and architectural features they offer. The presented architectures have been designed for a wide range of devices, supporting different security properties. Pieter Maene, Johannes Götzfried, Ruan de Clercq, Tilo Müller, Felix C. Freiling, Ingrid Verbauwhede |
IEEE Trans. Computers | 6 |
| 2018 | HEPCloud: An FPGA-Based Multicore Processor for FV Somewhat Homomorphic Function EvaluationabstractIn this paper, we present an FPGA based hardware accelerator ‘$\mathsf{HEPCloud}$’ for homomorphic evaluations of medium depth functions which has applications in cloud computing. Our$\mathsf{HEPCloud}$architecture supports the polynomial ring based homomorphic encryption scheme FV for a ring-LWE parameter set of dimension$2^{15}$, modulus size 1,228-bit, and a standard deviation 50. This parameter-set offers a multiplicative depth 36 and at least 85 bit security. The processor of$\mathsf{HEPCloud}$is composed of multiple parallel cores. To achieve fast computation time for such a large parameter-set, various optimizations in both algorithm and architecture levels are performed. For fast polynomial multiplications, we use CRT with NTT and achieve two dimensional parallelism in$\mathsf{HEPCloud}$. We optimize the BRAM access, use a fast Barrett like polynomial reduction method, optimize the cost of CRT, and design a fast divide-and-round unit. Beside parallel processing, we apply pipelining strategy in several of the sequential building blocks to reduce the impact of sequential computations. Finally, we implement$\mathsf{HEPCloud}$on a medium-size Xilinx Virtex 6 FPGA board ML605 board and measure its on-board performance. To store the ciphertexts during a homomorphic function evaluation, we use the large DDR3 memory of the ML605 board. Our FPGA-based implementation of$\mathsf{HEPCloud}$computes a homomorphic multiplication in 26.67 s, of which the actual computation takes only 3.36 s and the rest is spent for off-chip memory access. It requires about 37,551 s to evaluate the SIMON-64/128 block cipher, but the per-block timing is only about 18 s because$\mathsf{HEPCloud}$processes 2,048 blocks simultaneously. The results show that FPGA-based acceleration of homomorphic function evaluations is feasible, but fast memory interface is crucial for the performance. Sujoy Sinha Roy, Kimmo Järvinen 0001, Jo Vliegen, Frederik Vercauteren, Ingrid Verbauwhede |
IEEE Trans. Computers | 5 |
| 2018 | Private Mobile Pay-TV From Priced Oblivious TransferabstractIn pay-TV, a service provider offers TV programs and channels to users. To ensure that only authorized users gain access, conditional access systems (CAS) have been proposed. In existing CAS, users disclose to the service provider the TV programs and channels they purchase. We propose a pay-per-view and a pay-per-channel CAS that protect users' privacy. Our pay-per-view CAS employs priced oblivious transfer (POT) to allow a user to purchase TV programs without disclosing which programs were bought to the service provider. In our pay-per-channel CAS, POT is employed together with broadcast attribute-based encryption to achieve low storage overhead, collusion resistance, efficient revocation, and broadcast efficiency. We propose a new POT scheme and show its feasibility by implementing and testing our CAS on a representative mobile platform. Wouter Biesmans, Josep Balasch, Alfredo Rial, Bart Preneel, Ingrid Verbauwhede |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2017 | Fault Analysis of the ChaCha and Salsa Families of Stream Ciphers
Arthur Beckers, Benedikt Gierlichs, Ingrid Verbauwhede |
CARDIS | 3 |
| 2017 | SCM: Secure Code Memory ArchitectureabstractAn increasing number of applications implemented on a SoC (System-on-chip) require security features. This work addresses the issue of protecting the integrity of code and read-only data that is stored in memory. To this end, we propose a new architecture called SCM, which works as a standalone IP core in a SoC. To the best of our knowledge, there exists no architectural elements similar to SCM that offer the same strict security guarantees while, at the same time, not requiring any modifications to other IP cores in its SoC design. In addition, SCM has the flexibility to select the parts of the software to be protected, which eases the integration of our solution with existing software. The evaluation of SCM was done on the Zynq platform which features an ARM processor and an FPGA. The design was evaluated by executing a number of different benchmarks from memory protected by SCM, and we found that it introduces minimal overhead to the system. Ruan de Clercq, Ronald De Keulenaer, Pieter Maene, Bart Preneel, Bjorn De Sutter, Ingrid Verbauwhede |
AsiaCCS | 6 |
| 2017 | Fast Leakage Assessment
Oscar Reparaz, Benedikt Gierlichs, Ingrid Verbauwhede |
CHES | 3 |
| 2017 | Dude, is my code constant time?abstractThis paper introduces dudect: a tool to assess whether a piece of code runs in constant time or not on a given platform. We base our approach on leakage detection techniques, resulting in a very compact, easy to use and easy to maintain tool. Our methodology fits in around 300 lines of C and runs on the target platform. The approach is substantially different from previous solutions. Contrary to others, our solution requires no modeling of hardware behavior. Our solution can be used in black-box testing, yet benefits from implementation details if available. We show the effectiveness of our approach by detecting several variable-time cryptographic implementations. We place a prototype implementation of dudect in the public domain. Oscar Reparaz, Josep Balasch, Ingrid Verbauwhede |
DATE | 3 |
| 2017 | The Monte Carlo PUFabstractPhysically unclonable functions are used for IP protection, hardware authentication and supply chain security. While many PUF constructions have been put forward in the past decade, only few of them are applicable to FPGA platforms. Strict constraints on the placement and routing are the main disadvantages of the existing PUFs on FPGAs, because they place a high effort on the designer. In this paper we propose a new delay-based PUF construction called Monte Carlo PUF, that does not require low-level placement and routing control. This construction relies on the on-chip Monte Carlo method that is applied for measuring the delays of logic elements in order to extract a unique device fingerprint. The proposed construction allows a trade-off between the evaluation time and the error rate. The Monte Carlo PUF is implemented and evaluated on Xilinx Spartan-6 FPGAs. Vladimir Rozic, Bohan Yang 0001, Jo Vliegen, Nele Mentens, Ingrid Verbauwhede |
FPL | 5 |
| 2017 | SOFIA: Software and control flow integrity architecture
Ruan de Clercq, Johannes Götzfried, David Übler, Pieter Maene, Ingrid Verbauwhede |
Comput. Secur. | 5 |
| 2017 | Elliptic Curve Cryptography with Efficiently Computable Endomorphisms and Its Hardware Implementations for the Internet of ThingsabstractVerification of an ECDSA signature requires a double scalar multiplication on an elliptic curve. In this work, we study the computation of this operation on a twisted Edwards curve with an efficiently computable endomorphism, which allows reducing the number of point doublings by approximately 50 percent compared to a conventional implementation. In particular, we focus on a curve defined over the 207-bit prime field Fpwith p = 2207- 5,131. We develop several optimizations to the operation and we describe two hardware architectures for computing the operation. The first architecture is a small processor implemented in 0.13 μm CMOS ASIC and is useful in resource-constrained devices for the Internet of Things (IoT) applications. The second architecture is designed for fast signature verifications by using FPGA acceleration and can be used in the server-side of these applications. Our designs offer various trade-offs and optimizations between performance and resource requirements and they are valuable for IoT applications. Zhe Liu 0001, Johann Großschädl, Kimmo Järvinen 0001, Husen Wang, Ingrid Verbauwhede |
IEEE Trans. Computers | 6 |
| 2017 | Hardware Assisted Fully Homomorphic Function Evaluation and Encrypted SearchabstractIn this paper we propose a scheme to perform homomorphic evaluations of arbitrary depth with the assistance of a special module recryption box. Existing somewhat homomorphic encryption schemes can only perform homomorphic operations until the noise in the ciphertexts reaches a critical bound depending on the parameters of the homomorphic encryption scheme. The classical approach of bootstrapping also allows for arbitrary depth evaluations, but has a detrimental impact on the size of the parameters, making the whole setup inefficient. We describe two different instantiations of our recryption box for assisting homomorphic evaluations of arbitrary depth. The recryption box refreshes the ciphertexts by lowering the inherent noise and can be used with any instantiation of the parameters, i.e. there is no minimum size unlike bootstrapping. To demonstrate the practicality of the proposal, we design the recryption box on a Xilinx Virtex 6 FPGA board ML605 to support the FV somewhat homomorphic encryption scheme. The recryption box requires 0.43 ms to refresh one ciphertext. Further, we use this recryption box to boost the performance of encrypted search operation. On a 40 core Intel server, we can perform encrypted search in a table of 216 entries in around 20 seconds. This is roughly 20 times faster than the implementation without recryption box. Sujoy Sinha Roy, Frederik Vercauteren, Jo Vliegen, Ingrid Verbauwhede |
IEEE Trans. Computers | 4 |
| 2017 | LiBrA-CAN: Lightweight Broadcast Authentication for Controller Area NetworksabstractDespite realistic concerns, security is still absent from vehicular buses such as the widely used Controller Area Network (CAN). We design an efficient protocol based on efficient symmetric primitives, taking advantage of two innovative procedures: splitting keys between nodes and mixing authentication tags. This results in a higher security level when compromised nodes are in the minority, a realistic assumption for automotive networks. Experiments are performed on state-of-the-art Infineon TriCore controllers, contrasted with low-end Freescale S12X cores, while simulations are provided for the recently released CAN-FD standard. To gain compatibility with existent networks, we also discuss a solution based on CAN+. Bogdan Groza, Pal-Stefan Murvay, Anthony Van Herrewege, Ingrid Verbauwhede |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | High-Performance Ideal Lattice-Based Cryptography on 8-Bit AVR MicrocontrollersabstractOver recent years lattice-based cryptography has received much attention due to versatile average-case problems like Ring-LWE or Ring-SIS that appear to be intractable by quantum computers. In this work, we evaluate and compare implementations of Ring-LWE encryption and the bimodal lattice signature scheme (BLISS) on an 8-bit Atmel ATxmega128 microcontroller. Our implementation of Ring-LWE encryption provides comprehensive protection against timing side-channels and takes 24.9ms for encryption and 6.7ms for decryption. To compute a BLISS signature, our software takes 317ms and 86ms for verification. These results underline the feasibility of lattice-based cryptography on constrained devices. Zhe Liu 0001, Thomas Pöppelmann, Tobias Oder, Hwajeong Seo, Sujoy Sinha Roy, Tim Güneysu, Johann Großschädl, Howon Kim 0001, Ingrid Verbauwhede |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2017 | Sancus 2.0: A Low-Cost Security Architecture for IoT DevicesabstractThe Sancus security architecture for networked embedded devices was proposed in 2013 at the USENIX Security conference. It supports remote (even third-party) software installation on devices while maintaining strong security guarantees. More specifically, Sancus can remotely attest to a software provider that a specific software module is running uncompromised and can provide a secure communication channel between software modules and software providers. Software modules can securely maintain local state and can securely interact with other software modules that they choose to trust. Over the past three years, significant experience has been gained with applications of Sancus, and several extensions of the architecture have been investigated—both by the original designers as well as by independent researchers. Informed by these additional research results, this journal version of the Sancus paper describes an improved design and implementation, supporting additional security guarantees (such as confidential deployment) and a more efficient cryptographic core. We describe the design of Sancus 2.0 (without relying on any prior knowledge of Sancus) and develop and evaluate a prototype FPGA implementation. The prototype extends an MSP430 processor with hardware support for the memory access control and cryptographic functionality required to run Sancus. We report on our experience using Sancus in a variety of application scenarios and discuss some important avenues of ongoing and future work. Job Noorman, Jo Van Bulck, Jan Tobias Mühlberg, Frank Piessens, Pieter Maene, Bart Preneel, Ingrid Verbauwhede, Johannes Götzfried, Tilo Müller, Felix C. Freiling |
ACM Trans. Priv. Secur. | 7 |
| 2016 | Efficient Fuzzy Extraction of PUF-Induced Secrets: Theory and Applications
Jeroen Delvaux, Dawu Gu, Ingrid Verbauwhede, Matthias Hiller, Meng-Day (Mandel) Yu |
CHES | 3 |
| 2016 | On the Feasibility of Cryptography for a Wireless Insulin Pump SystemabstractThis paper analyses the security and privacy properties of a widely used insulin pump and its peripherals. We eavesdrop the wireless channel using Commercial Off-The-Shelf (COTS) software-based radios to intercept the messages sent between these devices; fully reverse-engineer the wireless communication protocol using a black-box approach; and document the message format and the protocol state-machine in use. The upshot is that no standard cryptographic mechanisms are applied and hence the system is shown to be completely vulnerable to replay and message injection attacks. Furthermore, sensitive patient health-related information is sent unencrypted over the wireless channel. Eduard Marin, Dave Singelée, Bohan Yang 0001, Ingrid Verbauwhede, Bart Preneel |
CODASPY | 4 |
| 2016 | TOTAL: TRNG on-the-fly testing for attack detection using Lightweight hardware
Bohan Yang 0001, Vladimir Rozic, Nele Mentens, Wim Dehaene, Ingrid Verbauwhede |
DATE | 5 |
| 2016 | SOFIA: Software and control flow integrity architecture
Ruan de Clercq, Ronald De Keulenaer, Bart Coppens 0001, Bohan Yang 0001, Pieter Maene, Koen De Bosschere, Bart Preneel, Bjorn De Sutter, Ingrid Verbauwhede |
DATE | 9 |
| 2016 | Software security: Vulnerabilities and countermeasures for two attacker models
Frank Piessens, Ingrid Verbauwhede |
DATE | 2 |
| 2016 | IoT: Source of test challengesabstractThe semiconductor industry has been driving a major part of its growth through first the PC and more recently the mobile market. Unfortunately, the PC market is in decline and also the end of the growth curve for mobile products is in sight now that virtually everyone on the planet has a smartphone and/or tablet. Hence, the semiconductor industry is putting its bets on `Internet of Things' (IoT) as the next application wave that will allow them to sell a lot of silicon real estate. Although what exactly IoT encompasses is under definition and hence still volatile, the first emerging products depict an image which is quite different from the traditional microprocessors or smartphone SOCs: small but with ubiquitous presence, wirelessly connected, energy harvesting, equipped with smart sensors, secure, and low cost. All these aspects have a profound impact on the challenges, solutions, and associated trade-offs for testing IoT chips and provide rich grounds for research. This paper provides seven views from different angles. Erik Jan Marinissen, Yervant Zorian, Mario Konijnenburg, Chih-Tsun Huang, Ping-Hsuan Hsieh, Peter Cockburn, Jeroen Delvaux, Vladimir Rozic, Bohan Yang 0001, Dave Singelée, Ingrid Verbauwhede, Cedric Mayor, Robert Van Rijsinge, Cocoy Reyes |
ETS | 11 |
| 2016 | Hardware acceleration of a software-based VPNabstractA Virtual Private Network (VPN) encrypts and decrypts the private traffic it tunnels over a public network. Maximizing the available bandwidth is an important requirement for network applications, but the cryptographic operations add significant computational load to VPN applications, limiting the network throughput. This work presents a coprocessor designed to offer hardware acceleration for these encryption and decryption operations. The open-source SigmaVPN application is used as the base solution, and a coprocessor is designed for the parts of Networking and Cryptography library (NaCl) which underlies the cryptographic operation of SigmaVPN. The hardware-software codesign of this work is implemented on a Xilinx Zynq-7000 SoC, showing a 93% reduction in the execution time of encrypting a 1024-byte frame, and this improved the TCP and UDP communication bandwidths by a factor of 4.36 and 5.36 respectively compared to pure software solution for a 1024-byte frame. Furkan Turan, Ruan de Clercq, Pieter Maene, Oscar Reparaz, Ingrid Verbauwhede |
FPL | 5 |
| 2016 | VLSI Design Methods for Low Power Embedded EncryptionabstractIntelligent things, medical devices, vehicles and factories, all part of cyberphysical systems, will only be secure if we can build devices that can perform the mathematically demanding cryptographic operations in an efficient way. Unfortunately, many of devices operate under extremely limited power, energy and area constraints. Yet we expect that they can execute, often in real-time, the symmetric key, public key and/or hash functions needed for the application. At the same time, we request that the implementations are also secure against a wide range of physical attacks. Ingrid Verbauwhede |
ACM Great Lakes Symposium on VLSI | 1 |
| 2016 | Binary decision diagram to design balanced secure logic stylesabstractEmbedded implementations of cryptographic algorithms require countermeasures against side-channel attacks (SCAs), that exploit physical variables measured during the computation. These countermeasures increase cost, power consumption and latency of the device. One class of countermeasures, hiding, consists of a balanced circuit style, including balancing of the capacitances and delays; it requires full connection to avoid memory effect that is an effect caused by repeatedly recharged energy after being only partially discharged at the internal parasitic capacitance. This paper proposes binary decision diagrams (BDDs) to derive complex pull-down networks that fulfill all these requirements while being compact at the same time; it uses sense amplifier-based logic (SABL) to obtain well-balanced pre-charge circuits. An attack based on mutual information analysis (MIA) is applied to the AES S-boxes implemented in our novel secure logic style. After the evaluation at pre-layout SPICE level, the balanced circuit with BDD leaks less information than comparable logic styles, even though the implementation area is reduced by 40.6%, the power consumption up to 46.1% and the delay by 35.2% compared to the classic SABL approach. Seokhie Hong, Bart Preneel, Ingrid Verbauwhede |
IOLTS | 4 |
| 2016 | Additively Homomorphic Ring-LWE Masking
Oscar Reparaz, Ruan de Clercq, Sujoy Sinha Roy, Frederik Vercauteren, Ingrid Verbauwhede |
PQCrypto | 5 |
| 2016 | Hold Your Breath, PRIMATEs Are Lightweight
Danilo Sijacic, Andreas B. Kidmose, Bohan Yang 0001, Subhadeep Banik, Begül Bilgin, Andrey Bogdanov, Ingrid Verbauwhede |
SAC | 7 |
| 2016 | Efficient Finite Field Multiplication for Isogeny Based Post Quantum Cryptography
Angshuman Karmakar, Sujoy Sinha Roy, Frederik Vercauteren, Ingrid Verbauwhede |
WAIFI | 4 |
| 2015 | Soteria: Offline Software Protection within Low-cost Embedded DevicesabstractProtecting the intellectual property of software that is distributed to third-party devices which are not under full control of the software author is difficult to achieve on commodity hardware today. Modern techniques of reverse engineering such as static and dynamic program analysis with system privileges are increasingly powerful, and despite possibilities of encryption, software eventually needs to be processed in clear by the CPU. To anyhow be able to protect software on these devices, a small part of the hardware must be considered trusted. In the past, general purpose trusted computing bases added to desktop computers resulted in costly and rather heavyweight solutions. In contrast, we present Soteria, a lightweight solution for low-cost embedded systems. At its heart, Soteria is a program-counter based memory access control extension for the TI MSP430 microprocessor. Based on our open implementation of Soteria as an openMSP430 extension, and our FPGA-based evaluation, we show that the proposed solution has a minimal performance, size and cost overhead while effectively protecting the confidentiality and integrity of an application's code against all kinds of software attacks including attacks from the system level. Johannes Götzfried, Tilo Müller, Ruan de Clercq, Pieter Maene, Felix C. Freiling, Ingrid Verbauwhede |
ACSAC | 6 |
| 2015 | Efficient Ring-LWE Encryption on 8-Bit AVR Processors
Zhe Liu 0001, Hwajeong Seo, Sujoy Sinha Roy, Johann Großschädl, Howon Kim 0001, Ingrid Verbauwhede |
CHES | 6 |
| 2015 | DPA, Bitslicing and Masking at 1 GHz
Josep Balasch, Benedikt Gierlichs, Oscar Reparaz, Ingrid Verbauwhede |
CHES | 4 |
| 2015 | A Masked Ring-LWE Implementation
Oscar Reparaz, Sujoy Sinha Roy, Frederik Vercauteren, Ingrid Verbauwhede |
CHES | 4 |
| 2015 | Lightweight Coprocessor for Koblitz Curves: 283-Bit ECC Including Scalar Conversion with only 4300 GatesabstractWe propose a lightweight coprocessor for 16-bit microcontrollers that implements high security elliptic curve cryptography. It uses a 283-bit Koblitz curve and offers 140-bit security. Koblitz curves offer fast point multiplications if the scalars are given as specific $$\tau $$ -adic expansions, which results in a need for conversions between integers and $$\tau $$ -adic expansions. We propose the first lightweight variant of the conversion algorithm and, by using it, introduce the first lightweight implementation of Koblitz curves that includes the scalar conversion. We also include countermeasures against side-channel attacks making the coprocessor the first lightweight coprocessor for Koblitz curves that includes a set of countermeasures against timing attacks, SPA, DPA and safe-error fault attacks. When the coprocessor is synthesized for 130 nm CMOS, it has an area of only 4,323 GE. When clocked at 16 MHz, it computes one 283-bit point multiplication in 98 ms with a power consumption of 97.70 $$\mu $$ W, thus, consuming 9.56 $$\mu $$ J of energy. Sujoy Sinha Roy, Kimmo Järvinen 0001, Ingrid Verbauwhede |
CHES | 3 |
| 2015 | Modular Hardware Architecture for Somewhat Homomorphic Function Evaluation
Sujoy Sinha Roy, Kimmo Järvinen 0001, Frederik Vercauteren, Vassil S. Dimitrov, Ingrid Verbauwhede |
CHES | 5 |
| 2015 | Consolidating Masking Schemes
Oscar Reparaz, Begül Bilgin, Svetla Nikova, Benedikt Gierlichs, Ingrid Verbauwhede |
CRYPTO (1) | 5 |
| 2015 | Highly efficient entropy extraction for true random number generators on FPGAsabstractTrue random number generators are essential components in cryptographic hardware. In this work, a novel entropy extraction method is used to improve throughput of jitter-based true random number generators on FPGA. By utilizing ultra-fast carry-logic primitives available on most commercial FPGAs, we have improved the efficiency of the entropy extraction, thereby increasing the throughput, while maintaining a compact implementation. Design steps and techniques are illustrated on an example of a ring-oscillator based true random number generator on Spartan-6 FPGA. In this design, the required accumulation time is reduced by 3 orders of magnitude compared to the most efficient oscillator-based TRNG on the same FPGA. The presented implementation occupies only 67 slices, achieves a throughput of 14.3 Mbps and it is provided with a formal evaluation of security. Vladimir Rozic, Bohan Yang 0001, Wim Dehaene, Ingrid Verbauwhede |
DAC | 4 |
| 2015 | Efficient software implementation of ring-LWE encryption
Ruan de Clercq, Sujoy Sinha Roy, Frederik Vercauteren, Ingrid Verbauwhede |
DATE | 4 |
| 2015 | Embedded HW/SW platform for on-the-fly testing of true random number generators
Bohan Yang 0001, Vladimir Rozic, Nele Mentens, Wim Dehaene, Ingrid Verbauwhede |
DATE | 5 |
| 2015 | On-the-fly tests for non-ideal true random number generatorsabstractHardware implementations of statistical tests are needed to detect failures and statistical weaknesses of entropy sources in True Random Number Generators on the fly. Current implementations of these tests work under the assumption that the entropy source produces independent, identically distributed (IID) numbers. However, some entropy sources produce non-IID data and rely on compression to provide the full entropy. Currently there are no embedded test implementations suitable for this type of entropy source. We provide the first FPGA implementation of embedded tests that estimate the generated min-Entropy and verify if it is within the expected boundaries. Bohan Yang 0001, Vladimir Rozic, Nele Mentens, Ingrid Verbauwhede |
ISCAS | 4 |
| 2015 | RECTANGLE: a bit-slice lightweight block cipher suitable for multiple platforms
Zhenzhen Bao, Dongdai Lin, Vincent Rijmen, Bohan Yang 0001, Ingrid Verbauwhede |
Sci. China Inf. Sci. | 6 |
| 2015 | Helper Data Algorithms for PUF-Based Key Generation: Overview and AnalysisabstractSecurity-critical products rely on the secrecy and integrity of their cryptographic keys. This is challenging for low-cost resource-constrained embedded devices, with an attacker having physical access to the integrated circuit (IC). Physically, unclonable functions are an emerging technology in this market. They extract bits from unavoidable IC manufacturing variations, remarkably analogous to unique human fingerprints. However, post-processing by helper data algorithms (HDAs) is indispensable to meet the stringent key requirements: reproducibility, high-entropy, and control. The novelty of this paper is threefold. We are the first to provide an in-depth and comprehensive literature overview on HDAs. Second, our analysis does expose new threats regarding helper data leakage and manipulation. Third, we identify several hiatuses/open problems in existing literature. Jeroen Delvaux, Dawu Gu, Dries Schellekens, Ingrid Verbauwhede |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | Anonymous Split E-Cash - Toward Mobile Anonymous PaymentsabstractAnonymous E-Cash was first introduced in 1982 as a digital, privacy-preserving alternative to physical cash. A lot of research has since then been devoted to extend and improve its properties, leading to the appearance of multiple schemes. Despite this progress, the practical feasibility of E-Cash systems is still today an open question. Payment tokens are typically portable hardware devices in smart card form, resource constrained due to their size, and therefore not suited to support largely complex protocols such as E-Cash. Migrating to more powerful mobile platforms, for instance, smartphones, seems a natural alternative. However, this implies moving computations from trusted and dedicated execution environments to generic multiapplication platforms, which may result in security vulnerabilities. In this work, we propose a new anonymous E-Cash system to overcome this limitation. Motivated by existing payment schemes based on MTM (Mobile Trusted Module) architectures, we consider at design time a model in which user payment tokens are composed of two modules: an untrusted but powerful execution platform (e.g., smartphone) and a trusted but constrained platform (e.g., secure element). We show how the protocol’s computational complexity can be relaxed by a secure split of computations: nonsensitive operations are delegated to the powerful platform, while sensitive computations are kept in a secure environment. We provide a full construction of our proposed Anonymous Split E-Cash scheme and show that it fully complies with the main properties of an ideal E-Cash system. Finally, we test its performance by implementing it on an Android smartphone equipped with a Java-Card-compatible secure element. Marijn Scheir, Josep Balasch, Alfredo Rial, Bart Preneel, Ingrid Verbauwhede |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2015 | Secure, Remote, Dynamic Reconfiguration of FPGAsabstractWith the widespread availability of broadband Internet, Field-Programmable Gate Arrays (FPGAs) can get remote updates in the field. This provides hardware and software updates, and enables issue solving and upgrade ability without device modification. In order to prevent an attacker from eavesdropping or manipulating the configuration data, security is a necessity. This work describes an architecture that allows the secure, remote reconfiguration of an FPGA. The architecture is partially dynamically reconfigurable and it consists of a static partition that handles the secure communication protocol and a single reconfigurable partition that holds the main application. Our solution distinguishes itself from existing work in two ways: it provides entity authentication and it avoids the use of a trusted third party. The former provides protection against active attackers on the communication channel, while the latter reduces the number of reliable entities. Additionally, this work provides basic countermeasures against simple power-oriented side-channel analysis attacks. The result is an implementation that is optimized toward minimal resource occupation. Because configuration updates occur infrequently, configuration speed is of minor importance with respect to area. A prototype of the proposed design is implemented, using 5,702 slices and having minimal downtime. Jo Vliegen, Nele Mentens, Ingrid Verbauwhede |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2015 | Accelerating Scalar Conversion for Koblitz Curve Cryptoprocessors on Hardware PlatformsabstractKoblitz curves are a class of computationally efficient elliptic curves where scalar multiplications can be accelerated using τNAF representations of scalars. However, conversion from an integer scalar to a short τNAF is a costly operation. In this paper, we improve the recently proposed scalar conversion scheme based on division by τ2. We apply two levels of optimizations in the scalar conversion architecture. First, we reduce the number of long integer subtractions during the scalar conversion. This optimization reduces the computation cost and also simplifies the critical paths present in the conversion architecture. Then we implement pipelines in the architecture. The pipeline splitting increases the operating frequency without increasing the number of cycles. We have provided detailed experimental results to support our claims made in this paper. Sujoy Sinha Roy, Junfeng Fan, Ingrid Verbauwhede |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Secure interrupts on low-end microcontrollersabstractEmbedded devices are increasingly becoming interconnected, sometimes over the public Internet. This poses a major security concern, as these devices handle sensitive information (e.g, banking credentials, personal data) or they are critical for the safety of human lives (e.g, smoke detector, airbag system). Security protocols need to be used in combination with a trusted computing base to ensure that attackers cannot alter the state of the software running on these devices to leak secrets. In this work we focus on the problem of secure interrupt handling, which has not been covered in related work. Our architecture for secure interrupts build on the idea of using simple memory isolation techniques to ensure leakage free processing of secret information on a microcontroller. Three methods of securely handling interrupts are proposed, each exploring a different tradeoff between hardware and software complexity, and interrupt latency. Prototype implementations based on an openMSP430 softcore demonstrate the practical feasibility of our architecture. Ruan de Clercq, Frank Piessens, Dries Schellekens, Ingrid Verbauwhede |
ASAP | 4 |
| 2014 | How to Use Koblitz Curves on Small Devices?
Kimmo Järvinen 0001, Ingrid Verbauwhede |
CARDIS | 2 |
| 2014 | Secure Lightweight Entity Authentication with Strong PUFs: Mission Impossible?
Jeroen Delvaux, Dawu Gu, Dries Schellekens, Ingrid Verbauwhede |
CHES | 4 |
| 2014 | Compact Ring-LWE Cryptoprocessor
Sujoy Sinha Roy, Frederik Vercauteren, Nele Mentens, Donald Donglong Chen, Ingrid Verbauwhede |
CHES | 5 |
| 2014 | Attacking PUF-Based Pattern Matching Key Generators via Helper Data Manipulation
Jeroen Delvaux, Ingrid Verbauwhede |
CT-RSA | 2 |
| 2014 | Ultra Low-Power implementation of ECC on the ARM Cortex-M0+abstractIn this work, elliptic curve cryptography (ECC) is used to make a fast, and very low-power software implementation of a public-key cryptography algorithm on the ARM Cortex-M0+. An optimization of the López-Dahab field multiplication method is proposed, which aims to reduce the number of memory accesses, as this is a slow operation on the target platform. A mixed C and assembly implementation was made; a random point multiplication requires 34.16 μJ, whereas our fixed point multiplication requires 20.63 μJ. Our implementation's energy consumption beats all other software implementations, on any platform, by a factor of at least 3.3. Ruan de Clercq, Leif Uhsadel, Anthony Van Herrewege, Ingrid Verbauwhede |
DAC | 4 |
| 2014 | Software Only, Extremely Compact, Keccak-based Secure PRNG on ARM Cortex-MabstractThe ability to generate secure random numbers is fundamental to the security of cryptographic protocols. Random Number Generators (RNGs) start to appear in recent modern Intel CPUs as used in desktops and servers. Solutions for embedded devices, such as e.g. sensor nodes and wireless routers, are still severely lacking however. Anthony Van Herrewege, Ingrid Verbauwhede |
DAC | 2 |
| 2014 | Key-recovery attacks on various RO PUF constructions via helper data manipulationabstractPhysically Unclonable Functions (PUFs) are security primitives that exploit the unique manufacturing variations of an integrated circuit (IC). They are mainly used to generate secret keys. Ring oscillator (RO) PUFs are among the most widely researched PUFs. In this work, we claim various RO PUF constructions to be vulnerable against manipulation of their public helper data. Partial/full key-recovery is a threat for the following constructions, in chronological order. (1) Temperature-aware cooperative RO PUFs, proposed at HOST 2009. (2) The sequential pairing algorithm, proposed at HOST 2010. (3) Group-based RO PUFs, proposed at DATE 2013. (4) Or more general, all entropy distiller constructions proposed at DAC 2013. Jeroen Delvaux, Ingrid Verbauwhede |
DATE | 2 |
| 2014 | Chaskey: An Efficient MAC Algorithm for 32-bit Microcontrollers
Nicky Mouha, Bart Mennink, Anthony Van Herrewege, Dai Watanabe, Bart Preneel, Ingrid Verbauwhede |
Selected Areas in Cryptography | 6 |
| 2014 | BLAKE-512-Based 128-Bit CCA2 Secure Timing Attack Resistant McEliece CryptoprocessorabstractThis paper presents a 128-bit CCA2-secure McEliece cryptoprocessor. The existing side-channel vulnerabilities in this regard are also taken care during the implementation of such a post-quantum immune code-based cryptosystem. In order to achieve CCA2 security on original McEliece algorithm, we incorporate a SHA-3 finalist, BLAKE-512 module into the architecture. A complete binary-XGCD algorithm for Goppa field is introduced. The final design on a Virtex-6 FPGA performs an encryption in${\bf 4.74}\nbsp \mu {\mbi{s}}$and a decryption in${\bf 0.92}\nbsp {\mbi{ms}}$. To the best of our knowledge, this is the first hardware design of McEliece with the above mentioned advanced security features which is also resistant against existing timing attacks. Santosh Ghosh, Ingrid Verbauwhede |
IEEE Trans. Computers | 2 |
| 2014 | Novel RNS Parameter Selection for Fast Modular MultiplicationabstractThe parameter selection of Residue Number Systems (RNS) has a great impact on its computational efficiency. This paper shows that a base extension, the most costly operation in RNS Montgomery multiplication, can be more efficient when the intervals between the RNS moduli are small. We propose a systematic RNS parameter selection procedure and two methods to select RNS moduli that lead to a reduced complexity. Our experimental results confirm the advantages of the selected moduli. Gavin Xiaoxu Yao, Junfeng Fan, Ray C. C. Cheung, Ingrid Verbauwhede |
IEEE Trans. Computers | 4 |
| 2013 | Inherent PUFs and secure PRNGs on commercial off-the-shelf microcontrollersabstractResearch on Physically Unclonable Functions (PUFs) has become very popular in recent years. However, all PUFs researched so far require either ASICs, FPGAs or a microcontroller with external components. Our research focuses on identifying PUFs in commercial off-the-shelf devices, e.g. microcontrollers. We show that PUFs exist in several off-theshelf products, which can be used for security applications. We present measurement results on the PUF behavior of five of the most popular microcontrollers today: ARM Cortex A,ARM Cortex-M,Atmel AVR, Microchip PIC16 and Texas Instruments MSP430. Based on these measurements, we can calculate whether these chips can be considered for applications requiring strong cryptography. As a result of these findings, we present a secure bootloader for the ARM Cortex-A9 platform based on a PUF inherent to the device, requiring no external components. Furthermore, instead of discarding the randomness in PUF responses, we utilize this to create strong seeds for pseudo-random number generators (PRNGs). The existence of a secure RNG is at the heart of virtually every cryptographic protocol, yet very often overlooked. We present the implementation of a strongly seeded PRNG for the ARM Cortex-M family, again requiring no external components. Anthony Van Herrewege, André Schaller, Stefan Katzenbeisser 0001, Ingrid Verbauwhede |
CCS | 4 |
| 2013 | On the Implementation of Unified Arithmetic on Binary Huff Curves
Santosh Ghosh, Amitabh Das, Ingrid Verbauwhede |
CHES | 4 |
| 2013 | A New Model for Error-Tolerant Side-Channel Cube Attacks
Zhenqi Li, Bin Zhang 0003, Junfeng Fan, Ingrid Verbauwhede |
CHES | 4 |
| 2013 | Low-energy encryption for medical devices: security adds an extra design dimensionabstractSmart medical devices will only be smart if they also include technology to provide security and privacy. In practice this means the inclusion of cryptographic algorithms of sufficient cryptographic strength. For battery operated devices or for passively powered devices, these cryptographic algorithms need highly efficient, low power, low energy realizations. Moreover, unique to cryptographic implementations is that they also need protection against physical tampering either active or passive. This means that countermeasures need to be included during the design process. Junfeng Fan, Oscar Reparaz, Vladimir Rozic, Ingrid Verbauwhede |
DAC | 4 |
| 2013 | High Precision Discrete Gaussian Sampling on FPGAs
Sujoy Sinha Roy, Frederik Vercauteren, Ingrid Verbauwhede |
Selected Areas in Cryptography | 3 |
| 2013 | Sancus: Low-cost Trustworthy Extensible Networked Devices with a Zero-software Trusted Computing Base
Job Noorman, Pieter Agten, Wilfried Daniels, Raoul Strackx, Anthony Van Herrewege, Christophe Huygens, Bart Preneel, Ingrid Verbauwhede, Frank Piessens |
USENIX Security Symposium | 8 |
| 2013 | Secure JTAG Implementation Using Schnorr Protocol
Amitabh Das, Jean DaRolt, Santosh Ghosh, Stefaan Seys, Sophie Dupuis, Giorgio Di Natale, Marie-Lise Flottes, Bruno Rouzeyre, Ingrid Verbauwhede |
J. Electron. Test. | 9 |
| 2013 | SPONGENT: The Design Space of Lightweight Cryptographic HashingabstractThe design of secure yet efficiently implementable cryptographic algorithms is a fundamental problem of cryptography. Lately, lightweight cryptography--optimizing the algorithms to fit the most constrained environments--has received a great deal of attention, the recent research being mainly focused on building block ciphers. As opposed to that, the design of lightweight hash functions is still far from being well investigated with only few proposals in the public domain. In this paper, we aim to address this gap by exploring the design space of lightweight hash functions based on the sponge construction instantiated with present-type permutations. The resulting family of hash functions is called spongent. We propose 13 spongent variants--or different levels of collision and (second) preimage resistance as well as for various implementation constraints. For each of them, we provide several ASIC hardware implementations--ranging from the lowest area to the highest throughput. We make efforts to address the fairness of comparison with other designs in the field by providing an exhaustive hardware evaluation on various technologies, including an open core library. We also prove essential differential properties of spongent permutations, give a security analysis in terms of collision and preimage resistance, as well as study in detail dedicated linear distinguishers. Andrey Bogdanov, Miroslav Knezevic, Gregor Leander, Deniz Toz, Kerem Varici, Ingrid Verbauwhede |
IEEE Trans. Computers | 6 |
| 2013 | Security Analysis of Industrial Test Compression SchemesabstractTest compression is widely used for reducing test time and cost of a very large scale integration circuit. It is also claimed to provide security against scan-based side-channel attacks. This paper pursues the legitimacy of this claim and presents scan attack vulnerabilities of test compression schemes used in commercial electronic design automation tools. A publicly available advanced encryption standard design is used and test compression structures provided by Synopsys, Cadence, and Mentor Graphics design for testability tools are inserted into the design. Experimental results of the differential scan attacks employed in this paper suggest that tools using X-masking and X-tolerance are vulnerable and leak information about the secret key. Differential scan attacks on these schemes have been demonstrated to have a best case success rate of 94.22% and 74.94%, respectively, for a random scan design. On the other hand, time compaction seems to be the strongest choice with the best case success rate of 3.55%. In addition, similar attacks are also performed on existing scan attack countermeasures proposed in the literature, thus experimentally evaluating their practical security. Finally, a suitable countermeasure is proposed and compared to the previously proposed countermeasures. Amitabh Das, Baris Ege, Santosh Ghosh, Lejla Batina, Ingrid Verbauwhede |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2013 | Hardware Designer's Guide to Fault AttacksabstractHardware designers invest a significant design effort when implementing computationally intensive cryptographic algorithms onto constrained embedded devices to match the computational demands of the algorithms with the stringent area, power, and energy budgets of the platforms. When it comes to designs that are employed in potential hostile environments, another challenge arises-the design has to be resistant against attacks based on the physical properties of the implementation, the so-called implementation attacks. This creates an extra design concern for a hardware designer. This paper gives an insight into the field of fault attacks and countermeasures to help the designer to protect the design against this type of implementation attacks. We analyze fault attacks from different aspects and expose the mechanisms they employ to reveal a secret parameter of a device. In addition, we classify the existing countermeasures and discuss their effectiveness and efficiency. The result of this paper is a guide for selecting a set of countermeasures, which provides a sufficient security level to meet the constraints of the embedded devices. Dusko Karaklajic, Jörn-Marc Schmidt, Ingrid Verbauwhede |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | A Speed Area Optimized Embedded Co-processor for McEliece CryptosystemabstractThis paper describes the systematic design methods of an embedded co-processor for a post quantum secure McEliece cryptosystem. A hardware/software co-design has been targeted for the realization of McEliece in practice on low-cost embedded platforms. Design optimizations take place when choosing system parameters, algorithm transformations, architecture choices, and arithmetic primitives. The final architecture consists of an 8-bit PicoBlaze softcore for flexibility and several parallel acceleration units for throughput optimization. A prototype of the co-processor is implemented on a Spartan-3an xc3s1400an FPGA, using less than 30% of its resources. On this FPGA, one McEliece decryption of an 80-bit security level takes less than 100K clock cycles corresponding to only 1 ms at a clock frequency of 92 MHz. This is 10 times faster and 3.8 times smaller than the existing design. Santosh Ghosh, Jeroen Delvaux, Leif Uhsadel, Ingrid Verbauwhede |
ASAP | 4 |
| 2012 | Interface Design for Mapping a Variety of RSA Exponentiation Algorithms on a HW/SW Co-design PlatformabstractWhen mapping public-key algorithms, such as RSA, onto constrained devices, both efficiency and flexibility are a challenge. Because word lengths are large, minimum 1024 bits, typically a dedicated co-processor is used. On the other hand, flexibility is required, because designers want to support a variety of RSA exponentiation algorithms. Typically the solution is then a hardware/software (HW/SW) co-design platform. In this paper we have chosen this approach: we use an 8051 micro-controller for flexibility and a Montgomery multiplier for efficiency. However, the importance of the interface between HW and SW is often neglected. The main focus of this paper is therefore to propose an interface that supports maximally the flexibility and the efficiency. We use this interface to compare six different exponentiation variants of RSA with and without side-channel attack countermeasures. Leif Uhsadel, Markus Ullrich, Ingrid Verbauwhede, Bart Preneel |
ASAP | 3 |
| 2012 | Theory and Practice of a Leakage Resilient Masking Scheme
Josep Balasch, Sebastian Faust, Benedikt Gierlichs, Ingrid Verbauwhede |
ASIACRYPT | 4 |
| 2012 | LiBrA-CAN: A Lightweight Broadcast Authentication Protocol for Controller Area Networks
Bogdan Groza, Pal-Stefan Murvay, Anthony Van Herrewege, Ingrid Verbauwhede |
CANS | 4 |
| 2012 | PUFs: Myth, Fact or Busted? A Security Evaluation of Physically Unclonable Functions (PUFs) Cast in Silicon
Stefan Katzenbeisser 0001, Ünal Koçabas, Vladimir Rozic, Ahmad-Reza Sadeghi, Ingrid Verbauwhede, Christian Wachsmann |
CHES | 5 |
| 2012 | PUFKY: A Fully Functional PUF-Based Cryptographic Key Generator
Roel Maes, Anthony Van Herrewege, Ingrid Verbauwhede |
CHES | 3 |
| 2012 | Selecting Time Samples for Multivariate DPA Attacks
Oscar Reparaz, Benedikt Gierlichs, Ingrid Verbauwhede |
CHES | 3 |
| 2012 | Power Analysis of Atmel CryptoMemory - Recovering Keys from Secure EEPROMs
Josep Balasch, Benedikt Gierlichs, Roel Verdult, Lejla Batina, Ingrid Verbauwhede |
CT-RSA | 5 |
| 2012 | PUF-based secure test wrapper design for cryptographic SoC testingabstractGlobalization of the semiconductor industry increases the vulnerability of integrated circuits. This particularly becomes a major concern for cryptographic IP blocks integrated on a System-on-Chip (SoC). The trustworthiness of these cryptographic blocks can be ensured with a secure test strategy. Presently, the IEEE 1500 Test Wrapper has emerged as the test standard for industrial SoCs. Additionally a secure activation mechanism has been proposed to this standard in order to restrict access to the testing interface to eligible testers by using a cryptographic authentication mechanism. This access mechanism is necessary in order not to provide any side-channels which may leak secret information for attackers. However, this approach requires the authentication mechanism to be implemented in hardware incurring an area overhead, and the authentication secrets to be securely stored in non-volatile memory (NVM), which may be susceptible to side-channel attacks. In this work, we enhance the secure test wrapper allowing testing of multiple IP blocks using a PUF-based authentication mechanism which overcomes the necessity of secure NVM and reduces the implementation overhead. Amitabh Das, Ünal Koçabas, Ahmad-Reza Sadeghi, Ingrid Verbauwhede |
DATE | 4 |
| 2012 | Low-cost implementations of on-the-fly tests for random number generatorsabstractRandom number generators (RNG) are important components in various cryptographic systems. Embedded security systems often require a high-quality digital source of randomness. Still, randomness of an RNG can vary due to aging effects, temperature or process conditions or intentional active attacks. This paper presents efficient, compact and reliable hardware implementations of 8 tests from the NIST test suite for statistical evaluation of randomness. These tests can be used for on-the-fly quality monitoring of on-chip random number generators as well as for fast hardware evaluation of RNG designs. Filip Veljkovic, Vladimir Rozic, Ingrid Verbauwhede |
DATE | 3 |
| 2012 | Differential Scan Attack on AES with X-tolerant and X-masked Test Response CompactorabstractScan-chains are test infrastructures included in a circuit for providing high fault coverage. However, they can be exploited by an attacker as a side-channel in the case of a cryptographic application like AES. Test Compression and thereafter X-tolerance and X-masking over it, which reduce test effort without compromising on testability, can help in counteracting scan-based attacks. This work focuses on the security issues of an AES-circuit containing test compression with X-masking and X-tolerance logic. With experimental results, we show the weakness of such an AES circuit against our modified differential scan-attack. Finally, the paper outlines two suitable countermeasures to prevent such attacks. Baris Ege, Amitabh Das, Santosh Ghosh, Ingrid Verbauwhede |
DSD | 4 |
| 2012 | Core Based Architecture to Speed Up Optimal Ate Pairing on FPGA Platform
Santosh Ghosh, Ingrid Verbauwhede, Dipanwita Roy Chowdhury |
Pairing | 2 |
| 2012 | Faster Pairing Coprocessor Architecture
Gavin Xiaoxu Yao, Junfeng Fan, Ray C. C. Cheung, Ingrid Verbauwhede |
Pairing | 4 |
| 2012 | A Practical Attack on KeeLoq
Wim Aerts, Eli Biham, Dieter De Moitie, Elke De Mulder, Orr Dunkelman, Sebastiaan Indesteege, Nathan Keller, Bart Preneel, Guy A. E. Vandenbosch, Ingrid Verbauwhede |
J. Cryptol. | 10 |
| 2012 | Extending ECC-based RFID authentication protocols to privacy-preserving multi-party grouping proofs
Lejla Batina, Yong Ki Lee, Stefaan Seys, Dave Singelée, Ingrid Verbauwhede |
Pers. Ubiquitous Comput. | 5 |
| 2012 | Efficient Hardware Implementation of Fp-Arithmetic for Pairing-Friendly CurvesabstractThis paper describes a new method to speed up {\hbox{\rlap{I}\kern 2.0pt{\hbox{F}}}}_p-arithmetic in hardware for pairing-friendly curves, such as the well-known Barreto-Naehrig (BN) curves. We explore the characteristics of the modulus defined by these curves and choose curve parameters such that {\hbox{\rlap{I}\kern 2.0pt{\hbox{F}}}}_p multiplication becomes more efficient. The proposed algorithm uses Montgomery reduction in a polynomial ring combined with a coefficient reduction phase using a pseudo-Mersenne number. As an application, we show that the performance of pairings on BN curves in hardware can be significantly improved, resulting in a factor 2.5 speedup compared with state-of-the-art hardware implementations. Junfeng Fan, Frederik Vercauteren, Ingrid Verbauwhede |
IEEE Trans. Computers | 3 |
| 2012 | A Pay-per-Use Licensing Scheme for Hardware IP Cores in Recent SRAM-Based FPGAsabstractCurrently achievable intellectual property (IP) protection solutions for field-programmable gate arrays (FPGAs) are limited to single large "monolithic" configurations. However, the ever growing capabilities of FPGAs and the consequential increasing complexity of their designs ask for a modular development model, where individual IP cores from multiple parties are integrated into a larger system. To enable such a model, the availability of IP protection at the modular level is imperative. In this work, we propose an IP protection mechanism for FPGA designs at the level of individual IP cores, by making use of the self-reconfiguring capabilities of modern FPGAs and deploying a trusted third party to run a metering service, similar to the work of Giineysu et ah and Drimer et at The proposed scheme makes it possible to enforce a pay-per-use licensing scheme which holds considerable advantages, both for IP core providers as well as for system integrators. Moreover, the scheme has a minimal implementation overhead and is the first of its kind to be solely based on primitives that are already available in recent commercially available FPGA devices. This allows for an immediate and feasible deployment, in contrast to earlier proposed solutions. Roel Maes, Dries Schellekens, Ingrid Verbauwhede |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Guest Editorial Integrated Circuit and System SecurityabstractThe Guest Editors' primary objective in organizing this Special Issue was to provide additional impetus for research in hardware security. Currently, the field is heavily dominated by testing, CAD, and IC researchers. They hope that researchers from other security fields will find the problems and the proposed solutions published here both interesting and important. Ten Special Issue papers are represented in this collection. Miodrag Potkonjak, Ramesh Karri, Ingrid Verbauwhede, Kouichi Itoh |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Fair and Consistent Hardware Evaluation of Fourteen Round Two SHA-3 CandidatesabstractThe first contribution of our paper is that we propose a platform, a design strategy, and evaluation criteria for a fair and consistent hardware evaluation of the second-round SHA-3 candidates. Using a SASEBO-GII field-programmable gate array (FPGA) board as a common platform, combined with well defined hardware and software interfaces, we compare all 256-bit version candidates with respect to area, throughput, latency, power, and energy consumption. Our approach defines a standard testing harness for SHA-3 candidates, including the interface specification for the SHA-3 module on our testing platform. The second contribution is that we provide both FPGA and 90-nm CMOS application-specific integrated circuit (ASIC) synthesis results and thereby are able to compare the results. Our third contribution is that we release the source code of all the candidates and by using a common, fixed, publicly available platform, our claimed results become reproducible and open for a public verification. Miroslav Knezevic, Kazuyuki Kobayashi, Jun Ikegami, Shin'ichiro Matsuo, Akashi Satoh, Ünal Koçabas, Junfeng Fan, Toshihiro Katashita, Takeshi Sugawara 0001, Kazuo Sakiyama, Ingrid Verbauwhede, Kazuo Ohta, Naofumi Homma, Takafumi Aoki |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2011 | spongent: A Lightweight Hash Function
Andrey Bogdanov, Miroslav Knezevic, Gregor Leander, Deniz Toz, Kerem Varici, Ingrid Verbauwhede |
CHES | 6 |
| 2011 | FPGA Implementation of Pairings Using Residue Number System and Lazy Reduction
Ray C. C. Cheung, Sylvain Duquesne, Junfeng Fan, Nicolas Guillermin, Ingrid Verbauwhede, Gavin Xiaoxu Yao |
CHES | 5 |
| 2011 | Low-cost fault detection method for ECC using Montgomery powering ladderabstractWhen using Elliptic Curve Cryptography (ECC) in constrained embedded devices such as RFID tags, López-Dahab's method along with the Montgomery powering ladder is considered as the most suitable method. It uses x-coordinate only for point representation, and meanwhile offers intrinsic protection against simple power analysis. This paper proposes a low-cost fault detection mechanism for Elliptic Curve Scalar Multiplication (ECSM) using the López-Dahab algorithm. Introducing minimal changes to the last round of the algorithm, we make it capable of detecting faults with a very high probability. In addition, by reusing the existing resources, we significantly reduce both performance losses and area overhead compared to other methods in this scenario. This method is suitable especially for constrained devices. Dusko Karaklajic, Junfeng Fan, Jörn-Marc Schmidt, Ingrid Verbauwhede |
DATE | 4 |
| 2011 | An In-depth and Black-box Characterization of the Effects of Clock Glitches on 8-bit MCUsabstractThe literature about fault analysis typically describes fault injection mechanisms, e.g. glitches and lasers, and cryptanalytic techniques to exploit faults based on some assumed fault model. Our work narrows the gap between both topics. We thoroughly analyse how clock glitches affect a commercial low-cost processor by performing a large number of experiments on five devices. We observe that the effects of fault injection on two-stage pipeline devices are more complex than commonly reported in the literature. While injecting a fault is relatively easy, injecting an exploitable fault is hard. We further observe that the easiest to inject and reliable fault is to replace instructions, and that random faults do not occur. Finally we explain how typical fault attacks can be mounted on this device, and describe a new attack for which the fault injection is easy and the cryptanalysis trivial. Josep Balasch, Benedikt Gierlichs, Ingrid Verbauwhede |
FDTC | 3 |
| 2011 | The Fault Attack Jungle - A Classification Model to Guide YouabstractFor a secure hardware designer, the vast array of fault attacks and countermeasures looks like a jungle. This paper aims at providing a guide through this jungle and at helping a designer of secure embedded devices to protect a design in the most efficient way. We classify the existing fault attacks on implementations of cryptographic algorithms on embedded devices according to different criteria. By doing do, we expose possible security threats caused by fault attacks and propose different classes of countermeasures capable of preventing them. Ingrid Verbauwhede, Dusko Karaklajic, Jörn-Marc Schmidt |
FDTC | 1 |
| 2011 | Physically unclonable functions: manufacturing variability as an unclonable device identifierabstractCMOS process variations are considered a burden to IC developers since they introduce undesirable random variability between equally designed ICs. However, it was demonstrated that measuring this variability can also be profitable as a physically unclonable method of silicon device identification. This can moreover be applied to generate strong cryptographic keys which are intrinsically bound to the embedding IC instance. This holds a number of very interesting advantages in comparison to traditional forms of secure identification and key storage. In this work, we summarize and compare the different proposed constructions and are able to identify some generalizing properties for PUFs on silicon devices. Ingrid Verbauwhede, Roel Maes |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | The cost of cryptography: Is low budget possible?abstractSummary form only given. Ambient intelligence, the future internet, smart dust, all lead to the immersion of electronics in the human environment. E-health applications are one example: patients will carry intelligent sensors and actuators which are wireless connected to monitoring devices and health professionals. All these applications carry heavy security and privacy risks. Strong authentication is needed such that the correct medical doses can be administered or that the settings of brain stimulants cannot be modified. On top, these devices have typically an extremely limited power, energy and area budget. So, the question is: can we provide secure implementations of cryptographic algorithms for these next generation applications? Cost can be expressed in memory footprint or gate count, number of clock cycles or time budget, power and energy budgets. On top, making implementations secure against physical attacks has an extra cost. In this presentation, we will discuss the implementation cost of several cryptographic algorithms, including public key, secret key and hash examples. We will indicate future directions in this field. Ingrid Verbauwhede |
IOLTS | 1 |
| 2011 | The communication and computation cost of wireless security: extended abstractabstract\n Contains fulltext :\n 92444.pdf (Publisher’s version ) (Open Access)\n Dave Singelée, Stefaan Seys, Lejla Batina, Ingrid Verbauwhede |
WISEC | 4 |
| 2011 | Design and design methods for unified multiplier and inverter and its application for HECC
Junfeng Fan, Lejla Batina, Ingrid Verbauwhede |
Integr. | 3 |
| 2011 | Tripartite modular multiplication
Kazuo Sakiyama, Miroslav Knezevic, Junfeng Fan, Bart Preneel, Ingrid Verbauwhede |
Integr. | 5 |
| 2010 | Implementation of binary edwards curves for very-constrained devicesabstractElliptic Curve Cryptography (ECC) is considered as the best candidate for Public-Key Cryptosystems (PKC) for ubiquitous security. Recently, Elliptic Curve Cryptography (ECC) based on Binary Edwards Curves (BEC) has been proposed and it shows several interesting properties, e.g., completeness and security against certain exceptional-points attacks. In this paper, we propose a hardware implementation of the BEC for extremely constrained devices. The w-coordinates and Montgomery powering ladder are used. Next, we also give techniques to reduce the register file size, which is the largest component of the embedded core. Thirdly, we apply gated clocking to reduce the overall power consumption. The implementation has a size of 13,427 Gate Equivalent (GE), and 149.5 ms are required for one point multiplication. To the best of our knowledge, this is the first hardware implementation of binary Edwards curves. Ünal Koçabas, Junfeng Fan, Ingrid Verbauwhede |
ASAP | 3 |
| 2010 | A compact FPGA-based architecture for elliptic curve cryptography over prime fieldsabstractThis paper proposes an FPGA-based application-specific elliptic curve processor over a prime field. This research targets applications for which compactness is more important than speed. To obtain a small datapath, the FPGA's dedicated multipliers and carry-chain logic are used and no parallellism is introduced. A small control unit is obtained by following a microcode approach, in which the instructions are stored in the FPGA's Block RAM. The use of algorithms that prevent Simple Power Analysis (SPA) attacks creates an extra cost in latency. Nevertheless, the created processor is flexible in the sense that it can handle all finite field operations over 256-bit prime fields and all elliptic curves of a specified form. The comparison with other implementations on the same generation of FPGAs learns that our design occupies the smallest area. Jo Vliegen, Nele Mentens, Jan Genoe, An Braeken, Serge Kubera, Abdellah Touhafi, Ingrid Verbauwhede |
ASAP | 7 |
| 2010 | Revisiting Higher-Order DPA Attacks:
Benedikt Gierlichs, Lejla Batina, Bart Preneel, Ingrid Verbauwhede |
CT-RSA | 4 |
| 2010 | An embedded platform for privacy-friendly road charging applicationsabstractSystems based on satellite localization are enabling new scenarios for road charging schemes by offering the possibility to charge drivers as a function of their road usage. An in-vehicle installation of a black box with the capabilities of a Location Based Service terminal suffices to deploy such a scheme. In the most straightforward architecture a back-end server collects vehicle's location data in order to extract the correct fees. However, with industry, governments and users being more and more aware of privacy issues the deployment of such system seems to be contradictory. Our contribution is the demonstration of a practical and functional road charging system based on PriPAYD [1]. Our black box is built guaranteeing most of the processing of location data in real-time, thus minimizing overheads required to ensure security and privacy. The performance of our software-based prototype is tested and proves that the deployment of a privacy-friendly solution can be achieved within a minimum cost increment compared to existing road charging schemes. Josep Balasch, Ingrid Verbauwhede, Bart Preneel |
DATE | 2 |
| 2010 | Low Cost Built in Self Test for Public Key Crypto CoresabstractThe testability of cryptographic cores brings an extra dimension to the process of digital circuits testing security. The benefits of the classical methods such as the scan-chain method introduce new vulnerabilities concerning the data protection. The Built-In Self-Test (BIST) is considered to be the most suitable countermeasure for this purpose. In this work we propose the use of a digit-serial multiplier over GF (2m), that is at the heart of many public-key cryptosystems, as a basic building block for the BIST circuitry. We show how the multiplier can be configuredto operate as a Test Pattern Generator and a Signature Analyzer. Furthermore, the multiplier becomes a fully self-testable design. All the additional features come at the cost of only a few extra gates. With a hardware overhead of 0.33 % this approach makes the multiplier perfectly suitable for low-end embedded devices. Dusko Karaklajic, Miroslav Knezevic, Ingrid Verbauwhede |
FDTC | 3 |
| 2010 | Breaking Elliptic Curve Cryptosystems Using Reconfigurable HardwareabstractThis paper reports a new speed record for FPGAs in cracking Elliptic Curve Cryptosystems. We conduct a detailed analysis of different F2(m)multiplication approaches in this application. A novel architecture using optimized normal basis multipliers is proposed to solve the Certicom challenge ECC2K-130. We compare the FPGA performance against CPUs, GPUs, and the Sony PlayStation 3. Our implementations show low-cost FPGAs outperform even multicore desktop processors and graphics cards by a factor of 2. Junfeng Fan, Daniel V. Bailey, Lejla Batina, Tim Güneysu, Christof Paar, Ingrid Verbauwhede |
FPL | 6 |
| 2010 | Privacy-Preserving ECC-Based Grouping Proofs for RFID
Lejla Batina, Yong Ki Lee, Stefaan Seys, Dave Singelée, Ingrid Verbauwhede |
ISC | 5 |
| 2010 | PrETP: Privacy-Preserving Electronic Toll Pricing
Josep Balasch, Alfredo Rial, Carmela Troncoso, Bart Preneel, Ingrid Verbauwhede, Christophe Geuens |
USENIX Security Symposium | 5 |
| 2010 | Speeding Up Bipartite Modular Multiplication
Miroslav Knezevic, Frederik Vercauteren, Ingrid Verbauwhede |
WAIFI | 3 |
| 2010 | Low-cost untraceable authentication protocols for RFIDabstractThe emergence of pervasive computing devices has raised several privacy issues. In this paper, we address the risk of tracking attacks in RFID networks. Our contribution is threefold: (1) We repair three revised EC-RAC protocols of Lee, Batina and Verbauwhede and show that two of the improved authentication protocols are wide-strong privacy-preserving and one wide-weak privacy-preserving; (2) We present the search protocol, a novel scheme which allows for privately querying a particular tag, and proof its security properties; and (3) We design a hardware architecture to demonstrate the implementation feasibility of our proposed solutions for a passive RFID tag. Due to the specific design of our authentication protocols, they can be realized with an area significantly smaller than other RFID schemes proposed in the literature, while still achieving the required security and privacy properties. Yong Ki Lee, Lejla Batina, Dave Singelée, Ingrid Verbauwhede |
WISEC | 4 |
| 2010 | Faster Interleaved Modular Multiplication Based on Barrett and Montgomery Reduction MethodsabstractIEEE Abstract—This paper proposes two improved interleaved modular multiplication algorithms based on Barrett and Montgomery modular reduction. The algorithms are simple and especially suitable for hardware implementations. Four large sets of moduli for which the proposed methods apply are given and analyzed from a security point of view. By considering state-of-the-art attacks on public-key cryptosystems, we show that the proposed sets are safe to use, in practice, for both elliptic curve cryptography and RSA cryptosystems. We propose a hardware architecture for the modular multiplier that is based on our methods. The results show that concerning the speed, our proposed architecture outperforms the modular multiplier based on standard modular multiplication by more than 50 percent. Additionally, our design consumes less area compared to the standard solutions. Index Terms—Modular multiplication, Barrett reduction, Montgomery reduction, public-key cryptography. Miroslav Knezevic, Frederik Vercauteren, Ingrid Verbauwhede |
IEEE Trans. Computers | 3 |
| 2009 | Faster -Arithmetic for Cryptographic Pairings on Barreto-Naehrig Curves
Junfeng Fan, Frederik Vercauteren, Ingrid Verbauwhede |
CHES | 3 |
| 2009 | Programmable and Parallel ECC Coprocessor Architecture: Tradeoffs between Area, Speed and Security
Xu Guo 0001, Junfeng Fan, Patrick Schaumont, Ingrid Verbauwhede |
CHES | 4 |
| 2009 | Low-Overhead Implementation of a Soft Decision Helper Data Algorithm for SRAM PUFs
Roel Maes, Pim Tuyls, Ingrid Verbauwhede |
CHES | 3 |
| 2009 | Case Study : A class E power amplifier for ISO-14443AabstractThis paper reports on the design and implementation of a class E push-pull amplifier in order to increase the reading range of an ISO-14443A RFID system. With the aid of classical design formulas and some alterations due to parasitic and intrinsic capacitances, a working implementation was made that can provide the loop with an amplified modulated current wave. Elke De Mulder, Wim Aerts, Bart Preneel, Ingrid Verbauwhede, Guy A. E. Vandenbosch |
DDECS | 4 |
| 2009 | Random numbers generation: Investigation of narrowtransitions suppression on FPGAabstractRandom number generators play an important role in the field of cryptography and security. It is often required that a random number generator consists of digital logic blocks only, so that it can be implemented on reconfigurable platforms. Since randomness cannot be proved by statistical tests there is a need for a provably secure hardware random number generator. In order to provide a proof of security, an experimental investigation of various physical effects on reconfigurable platforms is needed. In this paper we focus on the effect of narrow transitions suppression in the logic gates. The estimation of this effect may be crucial for the validity of the security proof of a RNG design. We explain our views on how experiments on FPGA should be performed and we give description of the measurement setup. We show that up to 98% of the transitions are suppressed in our experimental FPGA setup. Vladimir Rozic, Ingrid Verbauwhede |
FPL | 2 |
| 2009 | FPGA-based testing strategy for cryptographic chips: A case study on Elliptic Curve Processor for RFID tagsabstractTesting of cryptographic chips or components has one extra dimension: physical security. The chip designers should improve the design if it leaks too much information through side-channels, such as timing, power consumption, electric-magnetic radiation, and so on. This requires an evaluation of the security level of the chip under different side-channel attacks before it is manufactured. This paper presents an FPGA-based testing strategy for cryptographic chips. Using a block-based architecture, a testing bus and a shadow FPGA, we are able to check information leakage of each block. We describe this strategy with an Elliptic Curve Cryptosystem (ECC) for RFID tags. Junfeng Fan, Miroslav Knezevic, Dusko Karaklajic, Roel Maes, Vladimir Rozic, Lejla Batina, Ingrid Verbauwhede |
IOLTS | 7 |
| 2009 | Modular Reduction without Precomputational PhaseabstractIn this paper we show how modular reduction for integers with Barrett and Montgomery algorithms can be implemented efficiently without using a precomputational phase. We propose four distinct sets of moduli for which this method is applicable. The proposed modifications of existing algorithms are very suitable for fast software and hardware implementations of some public-key cryptosystems and in particular of Elliptic Curve Cryptography. Additionally, our results show substantial improvement when a small number of reductions with a single modulus is performed. Miroslav Knezevic, Lejla Batina, Ingrid Verbauwhede |
ISCAS | 3 |
| 2009 | A soft decision helper data algorithm for SRAM PUFsabstractIn this paper we propose the idea of using soft decision information in helper data algorithms (HDA). We derive and verify a distribution for the responses of SRAM-based physically unclonable functions (PUFs) and show that soft decision information becomes available without loss in min-entropy of the fuzzy secret. This significantly improves the implementation overhead of using an SRAM PUF + HDA for cryptographic key generation compared to previous constructions. Roel Maes, Pim Tuyls, Ingrid Verbauwhede |
ISIT | 3 |
| 2009 | Practical Mitigations for Timing-Based Side-Channel Attacks on Modern x86 ProcessorsabstractThis paper studies and evaluates the extent to which automated compiler techniques can defend against timing-based side-channel attacks on modern x86 processors. We study how modern x86 processors can leak timing information through side-channels that relate to control flow and data flow. To eliminate key-dependent control flow and key-dependent timing behavior related to control flow, we propose the use of if-conversion in a compiler backend, and evaluate a proof-of-concept prototype implementation. Furthermore, we demonstrate two ways in which programs that lack key-dependent control flow and key-dependent cache behavior can still leak timing information on modern x86 implementations such as the Intel Core 2 Duo, and propose defense mechanisms against them. Bart Coppens 0001, Ingrid Verbauwhede, Koen De Bosschere, Bjorn De Sutter |
SP | 2 |
| 2008 | Low-cost implementations of NTRU for pervasive securityabstractNTRU is a public-key cryptosystem based on the shortest vector problem in a lattice which is an alternative to RSA and ECC. This work presents a compact and low power NTRU design that is suitable for pervasive security applications such as RFIDs and sensor nodes. We have designed two architectures, one is only capable of encryption and the other one performs both encryption and decryption. The strategy for the designs includes clock gating of registers, operand isolation and precomputation. This work is also the first one to present a complete NTRU design with encryption/decryption circuitry. Our encryption-only NTRU design has a gate-count of 2:8 kgates and dynamic power consumption of 1:72μW. Moreover, encryption-decryption NTRU design consumes about 6μW dynamic power and consists of 10:5 kgates. Ali Can Atici, Lejla Batina, Junfeng Fan, Ingrid Verbauwhede, Siddika Berna Örs Yalçin |
ASAP | 4 |
| 2008 | On the high-throughput implementation of RIPEMD-160 hash algorithmabstractIn this paper we present two new architectures of the RIPEMD-160 hash algorithm for high throughput implementations. The first architecture achieves the iteration bound of RIPEMD-160, i.e. it achieves a theoretical upper bound on throughput at the micro-architecture level. The second architecture is designed by performing a gate level optimization and achieves a better performance than the first one at the cost of a larger gate area. Throughputs of 3.122 Gbps and 624 Mbps are achieved, with and without pipelining, respectively. Miroslav Knezevic, Kazuo Sakiyama, Yong Ki Lee, Ingrid Verbauwhede |
ASAP | 4 |
| 2008 | Power and Fault Analysis Resistance in Hardware through Dynamic Reconfiguration
Nele Mentens, Benedikt Gierlichs, Ingrid Verbauwhede |
CHES | 3 |
| 2008 | Fault Analysis Study of IDEA
Christophe Clavier, Benedikt Gierlichs, Ingrid Verbauwhede |
CT-RSA | 3 |
| 2008 | FPGA Design for Algebraic Tori-Based Public-Key CryptographyabstractAlgebraic torus-based cryptosystems are an alternative for Public-Key Cryptography (PKC). It maintains the security of a larger group while the actual computations are performed in a subgroup. Compared with RSA for the same security level, it allows faster exponentiation and much shorter bandwidth for the transmitted data. In this work we implement a torus-based cryptosystem, the so-called CEILIDH, on a multicore platform with an FPGA. This platform consists of a Xilinx MicroBlaze core and a multicore coprocessor. The platform supports CEILIDH, RSA and ECC over prime fields. The results show that one 170-bit torus T6exponentiation requires 20 ms, which is 5 times faster than 1024-bit RSA implementation on the same platform. Junfeng Fan, Lejla Batina, Kazuo Sakiyama, Ingrid Verbauwhede |
DATE | 4 |
| 2008 | Exploiting Hardware Performance CountersabstractWe introduce the usage of hardware performance counters (HPCs) as a new method that allows very precise access to known side channels and also allows access to many new side channels. Many current architectures provide hardware performance counters, which allow the profiling of software during runtime. Though they allow detailed profiling they are noisy by their very nature; HPC hardware is not validated along with the rest of the microprocessor. They are meant to serve as a relative measure and are most commonly used for profiling software projects or operating systems. Furthermore they are only accessible in restricted mode and can only be accessed by the operating system. We discuss this security model and we show first implementation results, which confirm that HPCs can be used to profile relatively short sequences of instructions with high precision. We focus on cache profiling and confirm our results by rerunning a recently published time based cache attack in which we replaced the time profiling function by HPCs. Leif Uhsadel, Andy Georges, Ingrid Verbauwhede |
FDTC | 3 |
| 2008 | Perfect Matching Disclosure Attacks
Carmela Troncoso, Benedikt Gierlichs, Bart Preneel, Ingrid Verbauwhede |
Privacy Enhancing Technologies | 4 |
| 2008 | Modular Reduction in GF(2n) without Pre-computational Phase
Miroslav Knezevic, Kazuo Sakiyama, Junfeng Fan, Ingrid Verbauwhede |
WAIFI | 4 |
| 2008 | A Cost-Effective Latency-Aware Memory Bus for Symmetric Multiprocessor SystemsabstractThis paper presents how a multi-core system can benefit from the use of a latency-aware memory bus capable of dual-concurrent data transfers on a single wire line: Source synchronous CDMA interconnect (SSCDMA-I) has been adopted to implement the memory bus of a shared-memory multi-core system. Two types of bus-based homogeneous and heterogeneous multi-core systems are modeled and simulated by a cycle-accurate simulation platform. Unlike the conventional time-division multiplexing (TDM) bus-based multi-core system that shows degradation in performance as the number of processing cores increases, the proposed SSCDMA bus-based multi-core shows higher performance up to 23.1% for 4 cores. The maximum latency of a heterogeneous multi-core system with a mix of traffic loads has been reduced up to 78%. These results demonstrate that the performance of multi-core systems can be improved with less cost and network complexity by reducing the bus contention interferences and by supporting higher concurrency in memory accesses that brings shorter critical word access latency. Jongsun Kim, Bo-Cheng Lai, Mau-Chung Frank Chang, Ingrid Verbauwhede |
IEEE Trans. Computers | 4 |
| 2008 | Elliptic-Curve-Based Security Processor for RFIDabstractRFID (Radio Frequency IDentification) tags need to include security functions, yet at the same time their resources are extremely limited. Moreover, to provide privacy, authentication and protection against tracking of RFID tags without loosing the system scalability, a public-key based approach is inevitable, which is shown by M. Burmester et al. In this paper, we present an architecture of a state-of-the-art processor for RFID tags with an Elliptic Curve (EC) processor over GF(2^163). It shows the plausibility of meeting both security and efficiency requirements even in a passive RFID tag. The proposed processor is able to perform EC scalar multiplications as well as general modular arithmetic (additions and multiplications) which are needed for the cryptographic protocols. As we work with large numbers, the register file is the most critical component in the architecture. By combining several techniques, we are able to reduce the number of registers from 9 to 6 resulting in EC processor of 10.1K gates. To obtain an efficient modulo arithmetic, we introduce a redundant modular operation. Moreover the proposed architecture can support multiple cryptographic protocols. The synthesis results with a 0.13 um CMOS technology show that the gate area of the most compact version is 12.5K gates. Yong Ki Lee, Kazuo Sakiyama, Lejla Batina, Ingrid Verbauwhede |
IEEE Trans. Computers | 4 |
| 2007 | Design methods for security and trustabstractThe design of ubiquitous and embedded computers focuses on cost factors such as area, power-consumption, and performance. Security and trust properties, on the other hand, are often an afterthought. Yet the purpose of ubiquitous electronics is to act and negotiate on their owner s behalf, and this makes trust a first-order concern. We outline a methodology for the design of secure and trusted electronic embedded systems, which builds on identifying the secure-sensitive part of a system (the root-of-trust) and iteratively partitioning and protecting that root-of-trust over all levels of design abstraction. This includes protocols, software, hardware, and circuits. We review active research in the area of secure design methodologies Ingrid Verbauwhede, Patrick Schaumont |
DATE | 1 |
| 2007 | Efficient pipelining for modular multiplication architectures in prime fieldsabstractThis paper presents a pipelined architecture of a modular Montgomery multiplier, which is suitable to be used in public key coprocessors. Starting from a baseline implementation of the Montgomery algorithm, a more compact pipelined version is derived. The design makes use of 16-bit integer multiplication blocks that are available on recently manufactured FPGAs. The critical path is optimized by omitting the exact computation of intermediate results in the Montgomery algorithm using a 6-2 carry-save notation. This results in a high-speed architecture,which outperforms previously designed Montgomery multipliers. Because a very popular application of Montgomery multiplication is public key cryptography, we compare our implementation to the state-of-the-art in Montgomery multipliers on the basis of performance results for 1024-bit RSA. Nele Mentens, Kazuo Sakiyama, Bart Preneel, Ingrid Verbauwhede |
ACM Great Lakes Symposium on VLSI | 4 |
| 2007 | Side-channel resistant system-level design flow for public-key cryptographyabstractIn this paper, we propose a new design methodology to assess the risk for side-channel attacks, more specifically timing analysis and simple power analysis, at an early design stage. This method is illustrated with the design of an elliptic curve cryptographic processor. It also allows to evaluate the quality of countermeasures against these attacks by evaluating hamming distances for eachsignal and each register in a partial functional domain (e.g. datapath or controller). Thus a first order side-channel-resistant design can be obtained with system-level design in which the simulation can run faster than conventional HDL simulations. Kazuo Sakiyama, Elke De Mulder, Bart Preneel, Ingrid Verbauwhede |
ACM Great Lakes Symposium on VLSI | 4 |
| 2007 | Secure IRIS VerificationabstractIn this paper, we present a novel secure iris verification system, where a transformed version of the iris template instead of the plain reference is stored for protecting the sensitive biometric data. An error correcting code (ECC) technique is adopted to perform the comparison in the transformed domain. A two-segment method is proposed to execute the feature verification, where a Bose-Chaudhuri-Hochquenghem (BCH) code of a random bit-stream is introduced to eliminate the considerable differences between the features extracted from different scans of irises. A reliable bits selection process during the iris feature generation stage reduces the system error rate from 6.0% to 0.8%. The appropriate size of the set of reliable bits is determined by investigating the best match between the associated error correct cutting edge and the actual verification accuracy. Shenglin Yang, Ingrid Verbauwhede |
ICASSP (2) | 2 |
| 2007 | Public-Key Cryptography on the Top of a NeedleabstractThis work describes the smallest known hardware implementation for Elliptic/Hyperelliptic Curve Cryptography (ECC/HECC). We propose two solutions for Public-key Cryptography (PKC), which are based on arithmetic on elliptic/hyperelliptic curves. One solution relies on ECC over binary fields 𝔽2𝓃where 𝓃 is a composite number of the form2𝑝(𝑝is a prime) and another on HECC on curves of genus 2 over 𝔽2𝑝. This implies the same arithmetic unit for both cases which supports arithmetic in a field 𝔽2𝑝. Our best solution that still results in a feasible performance features less than 5 kgates with an average power consumption smaller than 10μW. Lejla Batina, Nele Mentens, Kazuo Sakiyama, Bart Preneel, Ingrid Verbauwhede |
ISCAS | 5 |
| 2007 | HW/SW co-design of a hyperelliptic curve cryptosystem using a microcode instruction set coprocessor
Alireza Hodjat, Lejla Batina, David Hwang 0001, Ingrid Verbauwhede |
Integr. | 4 |
| 2007 | High-performance Public-key Cryptoprocessor for Wireless Mobile Applications
Kazuo Sakiyama, Lejla Batina, Bart Preneel, Ingrid Verbauwhede |
Mob. Networks Appl. | 4 |
| 2007 | Multicore Curve-Based Cryptoprocessor with Reconfigurable Modular Arithmetic Logic Units over GF(2n)abstractThis paper presents a reconfigurable curve-based cryptoprocessor that accelerates scalar multiplication of Elliptic Curve Cryptography (ECC) and HyperElliptic Curve Cryptography (HECC) of genus 2 over GF(2n). By allocating a copies of processing cores that embed reconfigurable Modular Arithmetic Logic Units (MALUs) over GF(2n), the scalar multiplication of ECC/HECC can be accelerated by exploiting Instruction-Level Parallelism (ILP). The supported field size can be arbitrary up to a(n + 1) - 1. The superscaling feature is facilitated by defining a single instruction that can be used for all field operations and point/divisor operations. In addition, the cryptoprocessor is fully programmable and it can handle various curve parameters and arbitrary irreducible polynomials. The cost, performance, and security trade-offs are thoroughly discussed for different hardware configurations and software programs. The synthesis results with a 0.13-mum CMOS technology show that the proposed reconfigurable cryptoprocessor runs at 292 MHz, whereas the field sizes can be supported up to 587 bits. The compact and fastest configuration of our design is also synthesized with a fixed field size and irreducible polynomial. The results show that the scalar multiplication of ECC over GF(2163) and HECC over GF(283) can be performed in 29 and 63 mus, respectively. Kazuo Sakiyama, Lejla Batina, Bart Preneel, Ingrid Verbauwhede |
IEEE Trans. Computers | 4 |
| 2007 | Design of an Interconnect Architecture and Signaling Technology for Parallelism in CommunicationabstractThe need for efficient interconnect architectures beyond the conventional time-division multiplexing (TDM) protocol-based interconnects has been brought on by the continued increase of required communication bandwidth and concurrency of small-scale digital systems. To improve the overall system performance without increasing communication resources and complexity, this paper presents a cost-effective interconnect architecture, communication protocol, and signaling technology that exploits parallelism in board-level communication, resulting in shorter latency and higher concurrency on a shared bus or link: the proposed source synchronous CDMA interconnect (SSCDMA-I) enables dual concurrent transactions on a single wire line as well as flexible input/output (I/O) reconfiguration. The SSCDMA-I utilizes 2-bit orthogonal CDMA coding and a variation of source synchronous clocking for multilevel superposition; a single 3-level SSCDMA-I line operates as if it consists of dual virtual time-multiplexed interconnects, which exploits communication parallelism with a reduced number of pins, wires, and complexity. The unique multiple access capability of the SSCDMA-I improves real-time communication between multiple semiconductor intellectual property (IP) blocks on a shared link or bus by reducing the bus contention interference from simultaneous traffic requests and by taking advantage of shorter request latency. The prototype transceiver chip is implemented in 0.18-m CMOS and the 10-cm test PC board system achieves an aggregate data rate of 2.5 Gb/s/pin between four off-chip (2Tx-to-2Rx) I/Os. Jongsun Kim, Ingrid Verbauwhede, Mau-Chung Frank Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | High Speed Channel Coding Architectures for the Uncoordinated OR ChannelabstractThough it promises high bandwidths, the optical medium is not popular in local area networks. This is because current optical networks do not offer the ease of use and setup that an uncoordinated multiple access network such as Ethernet offers. By careful design and implementation of high speed channel coding architectures, we show that it is possible for optical networks to exhibit these desirable properties while maintaining high optical transmission rates. This paper presents an interleaver-division multiple access (IDMA) architecture implemented with a rate 1/20, 64-state Viterbi decoder and a word-based interleaver. These structures allowed us to achieve optical data rates of 2Gbps in FPGA implementation and 5.4Gbps for 0.18mum ASIC implementation. The techniques presented can be adapted for other similar architectures Herwin Chan, Miguel Griot, Andres I. Vila Casado, Richard D. Wesel, Ingrid Verbauwhede |
ASAP | 5 |
| 2006 | Cross Layer Design to Multi-thread a Data-Pipelining Application on a Multi-processor on ChipabstractData-Pipelining is a widely used model to represent streaming applications. Incremental decomposition and optimization of a data-pipelining application onto a multi-processor platform spans multiple design layers, including the application layer, the system software layer, the architecture layer and the micro-architecture layer. For best results, designers have to consider multiple design layers (vertical exploration) and multiple architecture options (horizontal exploration). By using a data-pipelining JPEG encoder as the application driver, this paper presents a comprehensive analysis of mapping a data-pipelined application through multiple design layers, to a shared-memory SMP (Symmetric Multi- Processing) system. It is shown that a single-layered optimization ends up with a 110% worse design if the system effects from other layers are not taken into account. Compared to the nominal case, with appropriate mapping of the application, we achieve 47.5% improvement for high performance design and 77.6% energy reduction for energy efficient design under constant performance. Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede |
ASAP | 4 |
| 2006 | Throughput Optimized SHA-1 Architecture Using Unfolding TransformationabstractIn this paper, the authors analyze the theoretical delay bound of the SHA-1 algorithm and propose architectures to achieve high throughput hardware implementations which approach this bound. According to the results of FPGA implementations, 3,541 Mbps with a pipeline and 893 Mbps without a pipeline were achieved. Moreover, synthesis results using 0.18mum CMOS technology showed that 10.4 Gbps with a pipeline and 3.1 Gbps without a pipeline can be achieved. These results are much faster than previously published results. The high throughputs are due to the unfolding transformation, which reduces the number of required cycles for one block hash. The authors reduced the required number of cycles to 12 cycles for a 512 bit block and showed that 12 cycles is the optimal in our design Yong Ki Lee, Herwin Chan, Ingrid Verbauwhede |
ASAP | 3 |
| 2006 | Superscalar Coprocessor for High-Speed Curve-Based Cryptography
Kazuo Sakiyama, Lejla Batina, Bart Preneel, Ingrid Verbauwhede |
CHES | 4 |
| 2006 | Design with race-free hardware semanticsabstractMost hardware description languages do not enforce determinacy, meaning that they may yield races. Race conditions pose a problem for the implementation, verification, and validation of hardware. Enforcing determinacy at the modeling level provides a solution to this problem. In this paper, we consider a common model of computation for hardware modeling - a network of cycle-true finite-state-machines with datapaths (FSMDs) - and we identify the conditions under which such models are guaranteed to be race-free. We base our analysis on the Kahn principle and a formal framework to represent FSMD semantics. We present our conclusions as four simple and easy to enforce modeling rules. A hardware designer that applies those four modeling rules, will thus obtain race-free hardware Patrick Schaumont, Sandeep K. Shukla, Ingrid Verbauwhede |
DATE | 3 |
| 2006 | Reconfigurable Architectures for Curve-Based Cryptography on Embedded Micro-ControllersabstractThis paper discusses architectures for embedded security to enable various cryptographic services at low cost. To realize the large bit-lengths and complex arithmetic on an 8-bit embedded micro-controller, several hardware acceleration options for elliptic and hyperelliptic curve cryptography (ECC and HECC) are studied and systematically evaluated. Two key factors influence the performance: one is the communication interface i.e. I/O transfers between processor and co-processor and the other one is the boundary between hardware and software. Our experiments are run on an 8051 and an AVR micro-controller with the crypto co-processors implemented on a FPGA Lejla Batina, Alireza Hodjat, David Hwang 0001, Kazuo Sakiyama, Ingrid Verbauwhede |
FPL | 5 |
| 2006 | Fpga-Oriented Secure Data Path Design: Implementation of a Public Key CoprocessorabstractThis paper introduces a secure FPGA implementation of a coprocessor for public key cryptography. It supports Elliptic Curve Cryptography (ECC) as well as the older RSA standard. When choosing adequate key lengths, RSA and ECC are assumed to be secure from an algorithmic point of view. On the other hand, an implementation of these algorithms should also guarantee side-channel security. This feature does not only cause an inevitable performance degradation, but also an area increase. We overcome these drawbacks by fitting the public key architecture and algorithms into a coprocessor that optimally exploites the dedicated features on a Spartan XC3S4000. Although this is a very low-cost FPGA, the performance results of our implementation meet the requirements of a broad range of high-end applications. Nele Mentens, Kazuo Sakiyama, Lejla Batina, Ingrid Verbauwhede, Bart Preneel |
FPL | 4 |
| 2006 | FPGA Vendor Agnostic True Random Number GeneratorabstractThis paper describes a solution for the generation of true random numbers in a purely digital fashion; making it suitable for any FPGA type, because no FPGA vendor specific features (e.g., like phase-locked loop) or external analog components are required. Our solution is based on a framework for a provable secure true random number generator recently proposed by Sunar, Martin and Stinson. It uses a large amount of ring oscillators with identical ring lengths as a fast noise source - but with some deterministic bits - and eliminates the non-random samples by appropriate post-processing based on resilient functions. This results in a slower bit stream with high entropy. Our FPGA implementation achieves a random bit throughput of more than 2 Mbps, remains fairly compact (needing minimally 110 ring oscillators of 3 inverters) and is highly portable Dries Schellekens, Bart Preneel, Ingrid Verbauwhede |
FPL | 3 |
| 2006 | A Parallel Processing Hardware Architecture for Elliptic Curve CryptosystemsabstractWe propose a parallel processing crypto-processor for elliptic curve cryptography (ECC) to speed up EC point multiplication. The processor consists of a controller that dynamically checks instruction-level parallelism (ILP) and multiple sets of modular arithmetic logic units accelerating modular operations. A case study of HW design with the proposed architecture shows that EC point multiplication over GF(p) and GF(2m) can be improved by a factor of 1.6 compared to the case of using single processing element Kazuo Sakiyama, Elke De Mulder, Bart Preneel, Ingrid Verbauwhede |
ICASSP (3) | 4 |
| 2006 | Flexible hardware architectures for curve-based cryptographyabstractThis paper compares implementations of elliptic and hyperelliptic curve cryptography (ECC and HECC) on an FPGA platform. We use the same low-level blocks to implement the basic operations and we choose the bit-lengths so that both systems have equal security levels. The results are in favor of HECC. Our HECC implementation is slightly larger than ECC, but at the same time around 35% faster Lejla Batina, Nele Mentens, Bart Preneel, Ingrid Verbauwhede |
ISCAS | 4 |
| 2006 | A fast dual-field modular arithmetic logic unit and its hardware implementationabstractWe propose a fast modular arithmetic logic unit (MALU) that is scalable in the digit size (d) and the field size (k). The datapath of MALU has chains of carry save adders (CSAs) to speed up the large integer arithmetic operations over GF(p) and GF(2m). It is well suited and very efficient for the modular multiplication and addition/subtraction which are the computational kernels of elliptic curve and hyperelliptic curve cryptography (H/ECC). While maintaining the scalability and multi-function, we obtain a throughput of 205 Mbps and 388 Mbps with a clock rate of 110 MHz for 256-bit GF(p) and GF(2239) respectively on FPGA prototyping Kazuo Sakiyama, Bart Preneel, Ingrid Verbauwhede |
ISCAS | 3 |
| 2006 | Trellis Codes with Low Ones Density for the OR Multiple Access ChannelabstractThis paper presents trellis codes for the Z channel designed to maintain a relatively low ones density. These codes have applications in pulse-position modulation systems and as a solution for uncoordinated communication on the binary OR multiple-access channel (MAC). In this paper we consider the latter application to demonstrate the performance of the codes. The OR channel provides an unusual opportunity where single-user decoding permits operation at about 70% of the full multiple-access channel sum capacity. The interleaver-division multiple access technique applied in this paper should approach that performance with turbo solutions. However, the current paper focuses on very low latency codes with simple decoding, intended for very high speed (gigabits per second) applications. Namely, it focuses on nonlinear trellis codes that provide about 30% of the full multiple-access sum capacity at high speeds and with very low latency. These trellis codes are designed specifically for the Z-Channel that arises in a multiple-user OR channel, when the other users are treated as noise. In order to optimize the sum-capacity of the OR-MAC, the trellis code transmits codewords with a ones density much less than 50%. Also, a union bound technique that predicts the performance of these codes is presented. Results from simulations and a working FPGA implementation are shown. Miguel Griot, Andres I. Vila Casado, Wen-Yen Weng, Herwin Chan, Juthika Basak, Eli Yablonovitch, Ingrid Verbauwhede, Braham Jalali, Richard D. Wesel |
ISIT | 7 |
| 2006 | Area-Throughput Trade-Offs for Fully Pipelined 30 to 70 Gbits/s AES ProcessorsabstractThis paper explores the area-throughput trade-off for an ASIC implementation of the advanced encryption standard (AES). Different pipelined implementations of the AES algorithm as well as the design decisions and the area optimizations that lead to a low area and high throughput AES encryption processor are presented. With loop unrolling and outer-round pipelining techniques, throughputs of 30 Gbits/s to 70 Gbits/s are achievable in a 0.18-/spl mu/m CMOS technology. Moreover, by pipelining the composite field implementation of the byte substitution phase of the AES algorithm (inner-round pipelining), the area consumption is reduced up to 35 percent. By designing an offline key scheduling unit for the AES processor the area cost is further reduced by 28 percent, which results in a total reduction of 48 percent while the same throughput is maintained. Therefore, the over 30 Gbits/s, fully pipelined AES processor operating in the counter mode of operation can be used for the encryption of data on optical links. Alireza Hodjat, Ingrid Verbauwhede |
IEEE Trans. Computers | 2 |
| 2006 | Multilevel Design Validation in a Secure Embedded SystemabstractIn this paper, we present the simulation-based validation approach that we used during the design of ThumbPod-2, a portable fingerprint authentication system. The particular nature of secure system design has considerable impact on the simulation requirements and design flow. We present two key contributions. We will first show that rigorous design of secure digital systems requires a multilevel validation approach, meaning validation at multiple steps in the design flow. Indeed, an attacker chooses the easiest entry point and does not stick with one abstraction level. Second, we show the use of a cosimulation and codesign environment called GEZEL that can support this type of multilevel validation. We will illustrate this multilevel design validation strategy with the verification of security of the ThumbPod-2 device. Patrick Schaumont, David Hwang 0001, Shenglin Yang, Ingrid Verbauwhede |
IEEE Trans. Computers | 4 |
| 2006 | Clock-skew-optimization methodology for substrate-noise reduction with supply-current foldingabstractIn a synchronous clock distribution network with negligible skews, digital circuits switch simultaneously on the clock edge; therefore, they generate a lot of substrate noise due to the resulting sharp peaks on the supply current. A solution is to split a large design in different clock regions and introduce intentional clock skews between them, while taking the timing constraints into account. In this paper, the authors present a complete design flow to optimize the clock tree for less substrate-noise generation in large digital systems. It proposes a technique to assign combinatorial cells and flip-flops to the clock regions. It also takes into account the impact of unintentional clock skew such as jitter on the computed skews in order to assure a robust design. During the optimization, it uses compressed supply-current profiles to improve the CPU time. Experimental results show more than a factor-of-2 reduction in substrate-noise generation from large digital circuits of which the skews are optimized Mustafa Badaroglu, Kris Tiri, Geert Van der Plas, Piet Wambacq, Ingrid Verbauwhede, Stéphane Donnay, Georges Gielen, Hugo De Man |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2006 | A digital design flow for secure integrated circuitsabstractSmall embedded integrated circuits (ICs) such as smart cards are vulnerable to the so-called side-channel attacks (SCAs). The attacker can gain information by monitoring the power consumption, execution time, electromagnetic radiation, and other information leaked by the switching behavior of digital complementary metal-oxide-semiconductor (CMOS) gates. This paper presents a digital very large scale integrated (VLSI) design flow to create secure power-analysis-attack-resistant ICs. The design flow starts from a normal design in a hardware description language such as very-high-speed integrated circuit (VHSIC) hardware description language (VHDL) or Verilog and provides a direct path to an SCA-resistant layout. Instead of a full custom layout or an iterative design process with extensive simulations, a few key modifications are incorporated in a regular synchronous CMOS standard cell design flow. The basis for power analysis attack resistance is discussed. This paper describes how to adjust the library databases such that the regular single-ended static CMOS standard cells implement a dynamic and differential logic style and such that 20 000+ differential nets can be routed in parallel. This paper also explains how to modify the constraints and rules files for the synthesis, place, and differential route procedures. Measurement-based experimental results have demonstrated that the secure digital design flow is a functional technique to thwart side-channel power analysis. It successfully protects a prototype Advanced Encryption Standard (AES) IC fabricated in an 0.18-mum CMOS Kris Tiri, Ingrid Verbauwhede |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | An interactive codesign environment for domain-specific coprocessorsabstractEnergy-efficient embedded systems rely on domain-specific coprocessors for dedicated tasks such as baseband processing, video coding, or encryption. We present a language and design environment called GEZEL that can be used for the design, verification and implementation of such coprocessor-based systems.The GEZEL environment creates a platform simulator by combining a hardware simulation kernel with one or more instruction-set simulators. The hardware part of the platform is programmed in GEZEL, a deterministic, cycle-true and implementation-oriented hardware description language. GEZEL designs are scripted, allowing the hardware configuration of the platform simulator to be changed quickly without going through lengthy recompiles. For this reason, we call the environment interactive. We present the execution ladder as an optimization framework to balance interactivity against simulation speed.We demonstrate our approach using several designs including an AES encryption coprocessor and a Viterbi decoding coprocessor. We discuss the advantages of our approach as opposed to more conventional approaches using SystemC and Verilog/VHDL. Patrick Schaumont, Doris Ching, Ingrid Verbauwhede |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2005 | Side-channel aware design: Algorithms and Architectures for Elliptic Curve Cryptography over GF(2n)abstractThis paper proposes efficient algorithms for Elliptic Curve Cryptography (ECC). As an example a compact and efficient FPGA architecture for ECC over finite fields of even characteristic is presented. The implementation is balanced in order to increase the security w.r.t. simple side-channel attacks. Multiplication in GF(2 n ), Hardware implementation, Systolic array architecture, Elliptic Curve Cryptography (ECC), Montgomery method for point multiplication © 2005 IEEE. Lejla Batina, Nele Mentens, Bart Preneel, Ingrid Verbauwhede |
ASAP | 4 |
| 2005 | Hardware/Software Co-design for Hyperelliptic Curve Cryptography (HECC) on the 8051µP
Lejla Batina, David Hwang 0001, Alireza Hodjat, Bart Preneel, Ingrid Verbauwhede |
CHES | 5 |
| 2005 | Prototype IC with WDDL and Differential Routing - DPA Resistance Assessment
Kris Tiri, David Hwang 0001, Alireza Hodjat, Bo-Cheng Lai, Shenglin Yang, Patrick Schaumont, Ingrid Verbauwhede |
CHES | 7 |
| 2005 | A Systematic Evaluation of Compact Hardware Implementations for the Rijndael S-Box
Nele Mentens, Lejla Batina, Bart Preneel, Ingrid Verbauwhede |
CT-RSA | 4 |
| 2005 | Cooperative multithreading on 3mbedded multiprocessor architectures enables energy-scalable designabstractWe propose an embedded multiprocessor architecture and its associated thread-based programming model. Using a cycle-true simulation model of this architecture, we are able to estimate energy savings for a threaded C program. The savings are obtained by voltage- and frequency-scaling of the individual processors. We port a fingerprint minutiae detection application onto this architecture, and show the resulting performance on single-, dual-, and quad-processor configurations. The energy-scaled quadprocessor version results in a 77% energy reduction over the single-processor non-scaled implementation, at only a 2.2% degradation in cycle count. Patrick Schaumont, Bo-Cheng Lai, Ingrid Verbauwhede |
DAC | 4 |
| 2005 | A side-channel leakage free coprocessor IC in 0.18µm CMOS for embedded AES-based cryptographic and biometric processingabstractSecurity ICs are vulnerable to side-channel attacks (SCAs) that find the secret key by monitoring the power consumption and other information that is leaked by the switching behavior of digital CMOS gates. This paper describes a side-channel attack resistant coprocessor IC and its design techniques. The IC has been fabricated in 0.18µm CMOS. The coprocessor, which is used for embedded cryptographic and biometric processing, consists of four components: an Advanced Encryption Standard (AES) based cryptographic engine, a fingerprint-matching oracle, a template storage, and an interface unit. Two functionally identical coprocessors have been fabricated on the same die. The first, 'secure', coprocessor is implemented using a logic style called Wave Dynamic Digital Logic (WDDL) and a layout technique called differential routing. The second, 'insecure', coprocessor is implemented using regular standard cells and regular routing techniques. Measurement-based experimental results show that a differential power analysis (DPA) attack on the insecure coprocessor requires only 8,000 acquisitions to disclose the entire 128b secret key. The same attack on the secure coprocessor still does not disclose the entire secret key at 1,500,000 acquisitions. This improvement in DPA resistance of at least 2 orders of magnitude makes the attack de facto infeasible. The required number of measurements is larger than the lifetime of the secret key in most practical systems. Kris Tiri, David Hwang 0001, Alireza Hodjat, Bo-Cheng Lai, Shenglin Yang, Patrick Schaumont, Ingrid Verbauwhede |
DAC | 7 |
| 2005 | Simulation models for side-channel information leaksabstractSmall, embedded integrated circuits (ICs) such as smart cards are vulnerable to so-called side-channel attacks (SCAs). The attacker can gain information by monitoring the power consumption, execution time, electromagnetic radiation and other information that is leaked by the switching behavior of digital CMOS gates. Ever since power attacks have been introduced in 1999, many countermeasures have been proposed. Often a significant increase in security has been touted. We will show that in order to assess the effectiveness of a countermeasure, a correct simulation model of the side-channel information leaks is vital. We will show that seemingly correct approximations can lead to completely flawed results. Kris Tiri, Ingrid Verbauwhede |
DAC | 2 |
| 2005 | Design Method for Constant Power Consumption of Differential Logic CircuitsabstractSide channel attacks are a major security concern for smart cards and other embedded devices. They analyze the variations of the power consumption to find the secret key of the encryption algorithm implemented within the security IC. To address this issue, logic gates that have a constant power dissipation independent of the input signals are used in security ICs. The paper presents a design methodology to create fully connected differential pull down networks. Fully connected differential pull down networks are transistor networks that, for any complementary input combination, connect all the internal nodes of the network to one of the external nodes of the network. They are memoryless and, for that reason, have a constant load capacitance and power consumption. This type of network is used in specialized logic gates to guarantee a constant contribution of the internal nodes into the total power consumption of the logic gate. Kris Tiri, Ingrid Verbauwhede |
DATE | 2 |
| 2005 | A VLSI Design Flow for Secure Side-Channel Attack Resistant ICsabstractThe paper presents a digital VLSI design flow to create secure, side-channel attack (SCA) resistant integrated circuits. The design flow starts from a normal design in a hardware description language, such as VHDL or Verilog, and provides a direct path to an SCA resistant layout. Instead of a full custom layout or an iterative design process with extensive simulations, a few key modifications are incorporated in a regular synchronous CMOS standard cell design flow. We discuss the basis for side-channel attack resistance and adjust the library databases and constraints files of the synthesis and place-and-route procedures accordingly. Experimental results show that a DPA (differential power analysis) attack on a regular single ended CMOS standard cell implementation of a module of the DES algorithm discloses the secret key after 200 measurements. The same attack on a secure version still does not disclose the secret key after more than 2000 measurements. Kris Tiri, Ingrid Verbauwhede |
DATE | 2 |
| 2005 | Fast Dynamic Memory Integration in Co-Simulation Frameworks for Multiprocessor System on-ChipabstractThe paper proposes a technique to integrate and simulate a dynamic memory in a multiprocessor framework based on C/C++/SystemC. Using the host machine's memory management capabilities, dynamic data processing is supported without compromising speed and accuracy of the simulation. A first prototype in a shared memory context is presented. Oreste Villa, Patrick Schaumont, Ingrid Verbauwhede, Matteo Monchiero, Gianluca Palermo |
DATE | 3 |
| 2005 | A 3.84 gbits/s AES crypto coprocessor with modes of operation in a 0.18-µm CMOS technologyabstractIn this paper an AES crypto coprocessor that is fabricated using a 0.18-μm CMOS technology is presented. This crypto coprocessor performs the AES-128 encryption in both feedback and non-feedback modes of operation. A maximum throughput of 3.84 Gbits/s is achieved at a 330 MHz clock frequency for ECB, OFB, and CBC modes of operation. This crypto coprocessor can be programmed using the memory-mapped interface of an embedded CPU core and is tested using a LEON 32-bit (SPARC V8) processor in the ThumbPod secure system-on-chip. Alireza Hodjat, David Hwang 0001, Bo-Cheng Lai, Kris Tiri, Ingrid Verbauwhede |
ACM Great Lakes Symposium on VLSI | 5 |
| 2005 | Automatic secure fingerprint verification system based on fuzzy vault schemeabstractWe construct an automatic secure fingerprint verification system based on the fuzzy vault scheme to address a major security hole currently existing in most biometric authentication systems. The construction of the fuzzy vault during the enrollment phase is automated by aligning the most reliable reference points between different templates, based on which the converted features are used to form the lock set. The size of the fuzzy vault, the degree of the underlying polynomial, as well as the number of templates needed for reaching the reliable reference point are investigated. This results in a high unlocking complexity for attackers with an acceptable unlocking accuracy for legal users. Shenglin Yang, Ingrid Verbauwhede |
ICASSP (5) | 2 |
| 2005 | Energy and Performance Analysis of Mapping Parallel Multithreaded Tasks for An On-Chip Multi-Processor SystemabstractMultiprocessor systems offer superior performance and potentially better energy-reduction than single-processor systems. It all depends, however, on how well the application can be mapped onto the architecture. Indeed, a careful tradeoff of energy and performance requires a thorough understanding of the energy consumption pattern of the application across the architecture. We develop a simulation platform, MultiPo-Sim, which returns the cycle-accurate performance and energy consumption of a multiprocessor system, for both hardware components and software primitives. On the hardware level, energy scaling techniques can be modeled and each processing core can operate at different energy modes. MultiPo-Sim achieves 331K cycles per second simulation speed for a four-processor system on a 3GHz, 512MByte Fedora-2 PC. On the software level, data parallelizing and task parallelizing are two common models of multi-thread programming. By using MultiPo-Sim, we show that they show different energy and performance characteristics when mapping onto a multi-processor system. Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede |
ICCD | 4 |
| 2005 | Side-Channel Issues for Designing Secure Hardware ImplementationsabstractSelecting a strong cryptographic algorithm makes no sense if the information leaks out of the device through side-channels. Sensitive information, such as secret keys, can be obtained by observing the power consumption, the electromagnetic radiation, etc. This class of attacks is called side-channel attacks. Another type of attacks, namely fault attacks, reveal secret information by inserting faults into the device. Because both side-channel attacks and fault attacks are based on weaknesses in the implementation, they both belong to the category of implementation attacks. This work gives an overview of the state-of-the-art in implementation attacks, reviews the origin of this problem at the CMOS circuit level and discusses countermeasures. Lejla Batina, Nele Mentens, Ingrid Verbauwhede |
IOLTS | 3 |
| 2005 | Extended abstract: a race-free hardware modeling languageabstractWe describe race-free properties of a hardware description language called GEZEL. The language describes networks of cycle-true finite-state-machines with datapaths (FSMDs). We derive a set of four rules under which a network of such FSMDs satisfies the Kahn principle. When applying those rules, GEZEL programs will be determinate and a designer will thus obtain race-free hardware. We define extended FSMD networks as FSMD networks for which some components are user-defined and not specified as FSMDs. An important result is that the determinate properties of the FSMD network are also valid for the extended FSMD network provided that the user-defined components are determinate. Most hardware description languages do not have this determinacy. Their simulation semantics are dependent on simulator implementation, and on a run-time race resolution mechanism. We therefore position GEZEL as a model of computation that RTL designers should have in mind while creating RTL models. In fact, we can generate SystemC and other HDL code from GEZEL models, thereby guaranteeing the determinacy in the generated HDL code. Patrick Schaumont, Sandeep K. Shukla, Ingrid Verbauwhede |
MEMOCODE | 3 |
| 2005 | Platform-based design for an embedded-fingerprint-authentication deviceabstractFingerprint authentication, in an embedded and portable context, requires complex signal, network, and security-protocol processing in a resource-constrained implementation. We present a platform-based design approach for this application, based on a hierarchy of virtual machines (VM). The fingerprint authentication is programmed in Java, C, and VHSIC hardware description language, and mapped onto a hierarchy of three machines, consisting of an embedded Java VM, an Sparc-V8 core, and an field programmable gate array. We show how our approach is able to cope with multiple concurrent design processes and multiple application domains, including biometrics signal processing, as well as security-protocol implementation. The platform-based design approach also deals with reuse requirements for embedded software and hardware. The formulation of a platform as a VM enables design exploration and incremental design validation throughout the design traject, and results in a specialized, but still programmable, platform. The Java bytecode of our fingerprint authentication takes less than 10 kB. Patrick Schaumont, David Hwang 0001, Ingrid Verbauwhede |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Skiing the embedded systems mountainabstractUCLA teaches students how to master the steep slopes of the embedded systems mountain. The EE201A graduate course connects high-level design specification to embedded implementation. There is a long-standing and wide culture gap between system designers that create those abstract specifications and the system architects that need to implement them. In industry, the culture gap has separated software from hardware teams, and platform creators from platform users. In an embedded context, where these are very tightly connected, this leads to large inefficiencies both in design time and design results. Our course takes students to both sides of the gap and lets them look at this problem from different perspectives. For a given application, it teaches how to select target architectures, tools, and design methods. The course covers a stepwise systematic design process. It includes specification, transformation, and refinement of an application. Specifications enable systematic and structured expression of an application. Transformations rework specifications into ones that are a better match for a given target architecture. Refinements lower the abstraction level toward the target architecture. The embedded systems mountain is traversed in two directions. A vertical refinement axis covers elements such as power-memory-reduction methods or fixed-point refinement. A horizontal exploration axis covers various architecture alternatives including application-specific integrated circuits (ASIC), domain-specific processors, digital signal processors (DSP), embedded cores, programmable processors, and system-on-chip (SOC). During the course, the students also go through an extensive design project to apply the methods learned in this course. A typical embedded application is used to drive the project. In this paper it is illustrated using an embedded version of an image encoder, more specifically a JPEG encoder. Several commercial tools, design environments, and platforms have been used as alternative implementation targets for this application. Ingrid Verbauwhede, Patrick Schaumont |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2004 | Place and Route for Secure Standard Cell Design
Kris Tiri, Ingrid Verbauwhede |
CARDIS | 2 |
| 2004 | Java cryptography on KVM and its performance and security optimization using HW/SW co-design techniquesabstractThis paper describes a design approach to include and optimize Java based cryptographic applications into resource limited embedded devices. For easy prototyping and to be platform independent, the security applications are first developed in Java. Two Java cryptographic libraries, the Bouncy Castle API and the IAIK API are ported to a real embedded device for cost and performance evaluation. It requires 0.88Mbytes to 1.2Mbytes in the KVM footprint size and a few milliseconds to run secret key algorithms and message digests on a typical embedded device. In a second step, the performance critical components of the security applications are moved to hardware acceleration units. The GEZEL design environment is used for the hardware modeling and the co-simulation between software on KVM and the hardware co-processor. Moving the AES algorithm from the SH3-DSP microprocessor to a hardware co-processor shows a performance gain of 10.4x including the overhead in Java, C, and hardware interfaces. Then in a third step, the security critical components are realized by means of a special dynamic differential logic (DDL) style, which makes the secure modules resistant against side channel attacks. All key related actions and cryptographic algorithms are restricted to the secure co-processor. The overall performance gain is 25x compared to a pure Java implementation. Copyright 2004 ACM. Yusuke Matsuoka, Patrick Schaumont, Kris Tiri, Ingrid Verbauwhede |
CASES | 4 |
| 2004 | Interactive Cosimulation with Partial EvaluationabstractWe present a technique to improve the efficiency of hardware-software cosimulation, using design information known at simulator compile-time. The generic term for such optimization is partial evaluation. Our contribution is that we apply the optimization transparently to the user, and at multiple abstraction levels in the simulation. We use the technique to create an interactive codesign environment, and evaluate it on several designs including an AES encryption coprocessor and a Viterbi decoder, and for several instruction-set simulators. Compared to SystemC-based cosimulation, we achieve comparable cosimulation performance at only a fraction of the model-build time. Patrick Schaumont, Ingrid Verbauwhede |
DATE | 2 |
| 2004 | A Logic Level Design Methodology for a Secure DPA Resistant ASIC or FPGA ImplementationabstractThis paper describes a novel design methodology to implement a secure DPA resistant crypto processor. The methodology is suitable for integration in a common automated standard cell ASIC or FPGA design flow. The technique combines standard building blocks to make 'new' compound standard cells, which have a close to constant power consumption. Experimental results indicate a 50 times reduction in the power consumption fluctuations. Kris Tiri, Ingrid Verbauwhede |
DATE | 2 |
| 2004 | Architectures and Design Techniques for Energy Efficient Embedded DSP and Multimedia ProcessingabstractEnergy efficient embedded systems consist of a heterogeneous collection of very specific building blocks, connected together by a complex network of many dedicated busses and interconnect options. The trend to merge multiple functions into one device makes the design and integration of these "systems-on-chip" (SOC's) even more problematic. Yet, specifications and applications are never fixed and require the embedded units to be programmable. The topic of this paper is to give the designer architectures and design techniques to find the right balance between energy efficiency and flexibility. The key is to include programmability (or reconfiguration) at the right level of abstraction and tuned to the application domain. The challenge is to provide an exploration and programming environment for this heterogeneous architecture platform. Ingrid Verbauwhede, Patrick Schaumont, Christian Piguet, Bart Kienhuis |
DATE | 1 |
| 2004 | A 21.54 Gbits/s Fully Pipelined AES Processor on FPGAabstractThis paper presents the architecture of a fully pipelined AES encryption processor on a single chip FPGA. By using loop unrolling and inner-round and outer-round pipelining techniques, a maximum throughput of 21.54 Gbits/s is achieved. A fast and an area efficient composite field implementation of the byte substitution phase is designed using an optimum number of pipeline stages for FPGA implementation. A 21.54 Gbits/s throughput is achieved using 84 block RAMs and 5177 slices of a VirtexII-Pro FPGA with a latency of 31 cycles and throughput per area rate of 4.2 Mbps/Slice. Alireza Hodjat, Ingrid Verbauwhede |
FCCM | 2 |
| 2004 | Secure Logic Synthesis
Kris Tiri, Ingrid Verbauwhede |
FPL | 2 |
| 2004 | A realtime, memory efficient fingerprint verification systemabstractCreating a biometric verification system in an energy and area constrained embedded environment is a challenging problem. Our paper describes a secure and efficient embedded fingerprint verification system for the "ThumbPod" embedded device, in which a complete real-time fingerprint recognition module, including both the minutiae extraction and the matching, works on a 50 MHz LEON-2 processor. As a result of the proposed SW/HW accelerations and memory optimizations, we achieve 65% reduction on the execution time and 67% reduction on the memory size against the reference implementation. Shenglin Yang, Ingrid Verbauwhede |
ICASSP (5) | 2 |
| 2004 | Integrated Modeling and Generation of a Reconfigurable Network-on-ChipabstractSummary form only given. While a communication network is a critical component for an efficient system-on-chip multiprocessor, there are few approaches available to help with system-level architectural exploration of such a specialized interconnection network. We present an integrated modeling, simulation and implementation tool. A high level description of a network-on-chip can be simulated and converted into VHDL. The system simulation supports multiple instruction-set simulators, and obtains cycle-accurate performance metrics. This way, an optimal network configuration can be determined easily. We discuss our approach by designing a flexible network-on-chip and present implementation results after mapping into FPGA. The performance of our automatically generated network is comparable with a reference design directly developed in HDL. Doris Ching, Patrick Schaumont, Ingrid Verbauwhede |
IPDPS | 3 |
| 2004 | Embedded Software Integration for Coarse-Grain Reconfigurable SystemsabstractSummary form only given. Coarse-grain reconfigurable systems offer high performance and energy-efficiency, provided an efficient run-time reconfiguration mechanism is available. Using an embedded software vantage point, we define three levels of reconfigurability for such systems, each with a different degree of coupling between embedded software and reconfigurable hardware. We classify reconfigurable systems starting with tightly-coupled coprocessors and evolving to processor networks. This results in a gradual increase of energy-efficiency when compared to software-only systems, at the cost of increasing programming complexity. Using several sample applications including signal-, crypto-, and network-processing acceleration units, we demonstrate energy-efficiency improvements of 12 times over software for tightly-coupled systems up to 84 times for network-on-chip systems. Patrick Schaumont, Kazuo Sakiyama, Alireza Hodjat, Ingrid Verbauwhede |
IPDPS | 4 |
| 2004 | Reducing radio energy consumption of key management protocols for wireless sensor networksabstractThe security of sensor networks is a challenging area. Key management is one of the crucial parts in constructing the security among sensor nodes. However, key management protocols require a great deal of energy consumption, particularly in the transmission of initial key negotiation messages. In this paper, we examine three previously published sensor network security schemes: SPINS and C&R for master-key-based schemes, and Eschenhaur-Gligor (EG) for distributed-key-based schemes. We then present two new low-power schemes, which we call BROSK and OKS as alternatives to master-key-based schemes and distributed-key-based schemes, respectively. Compared to SPINS and C&R protocols, BROSK can reduce energy consumption by up to 12X by reducing the number of data transmissions in the key negotiation process. Compared with EG, OKS reduces energy by up to 96% and reduces memory requirements by up to 78%. Bo-Cheng Lai, David Hwang 0001, Sungha Pete Kim, Ingrid Verbauwhede |
ISLPED | 4 |
| 2003 | Finding the best system design flow for a high-speed JPEG encoderabstract26 students at the University of California, Los Angeles (UCLA) studied system level design methodologies through the design of a high-speed JPEG encoder. The results produced by 5 different design flows onto various target platforms demonstrate the high impact of tools on design quality. Kazuo Sakiyama, Patrick Schaumont, Ingrid Verbauwhede |
ASP-DAC | 3 |
| 2003 | Securing Encryption Algorithms against DPA at the Logic Level: Next Generation Smart Card Technology
Kris Tiri, Ingrid Verbauwhede |
CHES | 2 |
| 2003 | Design flow for HW / SW acceleration transparency in the thumbpod secure embedded systemabstractThis paper describes a case study and design flow of a secure embedded system called ThumbPod, which uses cryptographic and biometric signal processing acceleration. It presents the concept of HW/SW acceleration transparency, a systematic method to accelerate Java functions in both software and hardware. An example of acceleration transparency for a Rijndael encryption function is presented. The embedded prototype hardware platform is also described. Acceleration transparency yields software and hardware performance gains of 333X. David Hwang 0001, Bo-Cheng Lai, Patrick Schaumont, Kazuo Sakiyama, Shenglin Yang, Alireza Hodjat, Ingrid Verbauwhede |
DAC | 8 |
| 2002 | A Security Protocol for Biometric Smart Cards
David Hwang 0001, Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede |
CARDIS | 4 |
| 2002 | Clock tree optimization in synchronous CMOS digital circuits for substrate noise reduction using folding of supply current transientsabstractIn a synchronous clock distribution network with zero latencies, digital circuits switch simultaneously on the clock edge, therefore they generate substrate noise due to the sharp peaks on the supply current. We present a novel methodology optimizing the clock tree for less substrate generation by using statistical single cycle supply current profiles computed for every clock region taking the timing constraints into account. Our methodology is novel as it uses an error-driven compressed data set during the optimization over a number of clock regions specified for a significant reduction in substrate noise. It also produces a quality analysis of the computed latencies as a function of the clock skew. The experimental results show >x2 reduction of substrate noise generation from the circuits having four clock regions of which the latencies are optimized. Mustafa Badaroglu, Kris Tiri, Stéphane Donnay, Piet Wambacq, Hugo De Man, Ingrid Verbauwhede, Georges Gielen |
DAC | 6 |
| 2002 | Unlocking the design secrets of a 2.29 Gb/s Rijndael processorabstractThis contribution describes the design and performance testing of an Advanced Encryption Standard (AES) compliant encryption chip that delivers 2.29 GB/s of encryption throughput at 56 mw of power consumption. We discuss how the high level reference specification in C is translated into a parallel architecture. Design decisions are motivated from a system level viewpoint. The prototyping setup is discussed. Patrick Schaumont, Henry Kuo, Ingrid Verbauwhede |
DAC | 3 |
| 2002 | Guest editorial: low-power electronics and designabstractstatus: Published Enrico Macii, Ingrid Verbauwhede |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | Architectural Optimization for a 1.82Gbits/sec VLSI Implementation of the AES Rijndael Algorithm
Henry Kuo, Ingrid Verbauwhede |
CHES | 2 |
| 2001 | Panel: The Next HDL: If C++ is the Answer, What was the Question?abstractThe focus of this panel is on issues surrounding the use of C++ in modeling, integration of silicon IP and system-on-chip designs. In the last two years there have been several announcements promoting C++ based solutions and of multiple consortia (SystemC, Cynapps, Accellera, SpecC) that represent increasing commercial interest both from tool vendors as well as perhaps expression of genuine needs from the design houses. There are, however, serious questions about what value proposition does a C++ based design methodology bring to the IC or system designer? What has changed in the modeling technology (and/or available tools) that gives a new capability? Is synthesis the right target? or VAlidation? Tester modeling or testbench generation? This panel brings together advocates and opponents from the user community to highlight the achievements and the challenges that remain in use C++ for use in microelectronic circuits and systems. Rajesh K. Gupta 0001, Shishpal Rawat, Ingrid Verbauwhede, Gérard Berry, Ramesh Chandra, Daniel Gajski, Kris Konigsfeld, Patrick Schaumont |
DAC | 3 |
| 2001 | A Quick Safari Through the Reconfiguration JungleabstractCost effective systems use specialization to optimize factors such as power consumption, processing throughput, flexibility or combinations thereof. Reconfigurable systems obtain this specialization at rim-time. System reconfiguration has a vertical, a horizontal and a time dimension. We organize this design space as the reconfiguration hierarchy, and discuss the design methods that deal with it. Finally, we survey existing commercial platforms that support reconfiguration and situate them in the reconfiguration jungle. Patrick Schaumont, Ingrid Verbauwhede, Kurt Keutzer, Majid Sarrafzadeh |
DAC | 2 |
| 2001 | Low power showdown: comparison of five DSP platforms implementing an LPC speech codecabstractAn identical LPC speech coder has been implemented on a set of signal processing specific implementation platforms. The main goal of this experiment was to compare energy consumption. In addition, area/memory requirements and design time are also compared. The coder was first designed in floating-point C. Then, the fixed-point wordlengths were determined. Depending on the platform, either compiled code was generated, assembly code written or a Verilog/VHDL design was created. The platforms reported in this paper include the DSP processors TI C55/spl times/, TI C54/spl times/, TI C6/spl times/ and the design environments Ocapi and A|RT Designer. Energy consumption ranged from 2 /spl mu/J to 288 /spl mu/J per speech frame. Upon scaling the results to the same technology, our results indicated that the lowest power DSP processor (TI C55/spl times/) still consumes a factor of four more energy than an application specific processor. David Hwang 0001, Cimarron Mittelsteadt, Ingrid Verbauwhede |
ICASSP | 3 |
| 2000 | Low power DSP's for wireless communications (embedded tutorial session)abstractWireless communications and more specifically, the fast growing penetration of cellular phones and cellular infrastructure are the major drivers for the development of new programmable Digital Signal Processors (DSPs). In this tutorial, an overview will be given of recent developments in DSP processor architectures, that makes them well suited to execute computationally intensive algorithms typically found in communications systems. DSP processors have adapted instruction sets, memory architectures and data paths to execute compute intensive communications algorithms efficiently and in a low power fashion. Basic building blocks include convolutional decoders (mainly the Viterbi algorithm), turbo coding algorithms, FIR filters, speech coders, etc. This is illustrated with examples of different commercial and research processors. Please note that the authors do not endorse the processors used in this tutorial. These processors are used to illustrate how different solutions are proposed for the same problem. Ingrid Verbauwhede, Chris J. Nicol |
ISLPED | 1 |
| 1994 | Memory Estimation for High Level SynthesisabstractArticle Memory estimation for high level synthesis Share on Authors: Ingrid M. Verbauwhede EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CA EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CAView Profile , Chris J. Scheers Zycad Corporation, Fremont, CA Zycad Corporation, Fremont, CAView Profile , Jan M. Rabaey EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CA EECS Department, University of California at Berkeley, Cory Hall, Berkeley, CAView Profile Authors Info & Claims DAC '94: Proceedings of the 31st annual Design Automation ConferenceJune 1994 Pages 143–148https://doi.org/10.1145/196244.196313Online:06 June 1994Publication History 55citation265DownloadsMetricsTotal Citations55Total Downloads265Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ingrid Verbauwhede, Chris J. Scheers, Jan M. Rabaey |
DAC | 1 |
| 1994 | Specification and support for multidimensional DSP in the SILAGE languageabstractData flow languages are a natural way to describe the flow of computations in a DSP application. The SILAGE language has been developed for this purpose. It contains also more-dimensional arrays of signals and a natural extension of it, delayed versions of arrays, e.g. to represent previous frames in video applications. The paper describes new data flow analysis techniques, to support multi-dimensional arrays. It checks single assignment of arrays, checks if for each consumption of an indexed signal, there is a production, and it will create data dependencies between productions and consumptions. These problems are formulated as integer linear programming problems. This formulation is independent of the number of signals in the arrays. Results show very fast running times (> Ingrid Verbauwhede, Chris J. Scheers, Jan M. Rabaey |
ICASSP (2) | 1 |