Tim Güneysu

dblp:50/6307 · also Tim Erhan Güneysu · DBLP profile ↗
← Back
109ranked-venue papers
16as first author
34since 2021 · last 2026
0000-0002-3293-4989ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 55 · 5 first-author · 22 since 2021Systems, architecture and hardware · 53 · 11 first-author · 12 since 2021Software engineering, systems software and programming languages · 11 · 6 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 PaCMan - Partition-Code Masking for Combined Security
Fabian Buschkowski, Jakob Feldtkeller, Tim Güneysu, Elisabeth Krahmer, Jan Richter-Brockmann, Pascal Sasdrich
EUROCRYPT (7)3
2025 Revisiting Prime+Prune+Probe: Pitfalls and Remedies
abstract
Randomizing the mapping of memory addresses to cache locations is a promising approach for protecting computer systems against cache attacks. Multiple randomized caches have been proposed recently, with the aim of preventing adversaries from creating eviction sets - collections of addresses that compete with target memory addresses on cache space. However, Purnal et al. (IEEE SP 2021) demonstrated the Prime+ Prune+ Probeattack, which allows attackers to efficiently build generalized eviction sets, enabling target memory address eviction with a high probability. As the complexity of constructing eviction set is a key factor in randomized cache design, the Prime+prune+probe attack significantly reduces the security bounds of these randomizing designs. Since the Prime+prune+probe attack is probabilistic, generalized eviction sets often get stuck after repeated use, making them ineffective for typical cache attack settings. Prior works have noticed this behavior and proposed mitigation approaches. These approaches are based on evicting members of the eviction set from the cache, either probabilistically, using random memory accesses, or directly, using dedicated flush instructions. However, these proposals do not analyze the effectiveness of the techniques or evaluate their success. In this work we revisit Prime+prune+probe and analyze it in light of the possibility of eviction sets getting stuck. We observe that flushing does not behave as anticipated in realistic cache architectures, where invalid cache lines are filled first before evicting other lines. We further propose a new technique for allowing repeated attacks - combining random noise with flushing. We conduct an in-depth analysis of all discussed techniques and compare their complexity attacking an AES T-table implementation. We find that combining probabilistic eviction with flushing outperforms the traditional approaches by a factor of two, allowing attackers to increase the granularity and observe victim processes even better than in prior works.
Moritz Peters, Florian Stolz, Jan Philipp Thoma, Tim Güneysu, Yuval Yarom
ACSAC4
2025 To Extend or Not to Extend: Agile Masking Instructions for PQC
Markus Krausz, Georg Land, Florian Stolz, Jan Richter-Brockmann, Tim Güneysu
CANS5
2025 Multi-Partner Project: Securing Future Edge-AI Processors in Practice (CONVOLVE)
abstract
Artificial Intelligence (AI) has had a profound impact on our contemporary society, and it is indisputable that it will continue to play a significant role in the future. To further enhance AI experience and performance, a transition from large-scale server applications towards AI-powered edge devices is inevitable. In fact, current projections indicate that the market for Smart Edge Processors (SEPs) will grow beyond 70 Billion USD by 2026 [1]. Such a shift comes with major challenges, as these devices have limited computing and energy resources yet need to be highly performant. Additionally, security mechanisms need to be implemented to protect against diverse attack vectors as attackers now have physical access to the device. Besides cryptographic keys, Intellectual Property (IP), including neural network weights, may also be potential targets. The CONVOLVE [2] project (currently in its intermediate stage) follows a holistic approach to address these challenges and establish the EU in a leading position in embedded, ultra-low-power and secure processors for edge computing. It encompasses novel hardware technologies, end-to-end integrated workflows, and a security-by-design approach. This paper highlights the security aspects of future edge-AI processors by illustrating challenges encountered in CONVOLVE, the solutions we pursue including some early results, and directions for future research.
Sven Argo, Henk Corporaal, Alejandro Garza, Marc Geilen, Manil Dev Gomony, Tim Güneysu, Adrian Marotzke, Fouwad Jamil Mir, Jan Richter-Brockmann, Jeffrey Smith 0001, Mottaqiallah Taouil, Said Hamdioui
DATE6
2025 INDIANA - Verifying (Random) Probing Security Through Indistinguishability Analysis
Christof Beierle, Jakob Feldtkeller, Anna Guinet, Tim Güneysu, Gregor Leander, Jan Richter-Brockmann, Pascal Sasdrich
EUROCRYPT (8)4
2025 Multivariate TVLA - Efficient Side-Channel Evaluation Using Confidence Intervals
abstract
Securing cryptographic hardware designs and software implementations against side-channel attacks that leverage the power consumption or electromagnetic emanations of a device is an active topic of research. Different countermeasures against these attacks have been published, many of which rely on masking where sensitive information is split into multiple shares. Here, the information is hidden in higher statistical moments of the leakage if processed at the same time (univariate) or in combinations of side-channel information from different points in time (multivariate) if processed sequentially. Test Vector Leakage Assessment (TVLA) is a common evaluation technique to address the growing number of specific attacks. However, the assessment of multivariate leakage requires the evaluation of all possible combinations of sample points, massively slowing down the evaluation and in turn the development of countermeasures due to computational complexity.In this work, we develop and compare techniques to determine clock cycle combinations that leak information in a multivariate setting. We develop an efficient multivariate assessment framework and show how this approach can be used to generate evaluation results that satisfy a desired confidence level. Eventually, we demonstrate the practical relevance of our approach by applying it to two masked implementations of block ciphers.
Florian Bache, Jonas Wloka, Pascal Sasdrich, Tim Güneysu
IEEE Trans. Computers4
2024 On The Effect of Replacement Policies on The Security of Randomized Cache Architectures
abstract
Randomizing the mapping of addresses to cache entries has proven to be an effective technique for hardening caches against contention-based attacks like Prime+Probe. While attacks and defenses are still evolving, it is clear that randomized caches significantly increase the security against such attacks. However, one aspect that is missing from most analyses of randomized cache architectures is the choice of the replacement policy. Often, only the random- and LRU replacement policies are investigated. However, LRU is not applicable to randomized caches due to its immense hardware overhead, while the random replacement policy is not ideal from a performance and security perspective.
Moritz Peters, Nicolas Gaudin, Jan Philipp Thoma, Vianney Lapotre, Pascal Cotret, Guy Gogniat, Tim Güneysu
AsiaCCS7
2024 Formal Definition and Verification for Combined Random Fault and Random Probing Security
Sonia Belaïd, Jakob Feldtkeller, Tim Güneysu, Anna Guinet, Jan Richter-Brockmann, Matthieu Rivain, Pascal Sasdrich, Abdul Rahman Taleb
ASIACRYPT (7)3
2024 Practical Post-Quantum Signatures for Privacy
abstract
The transition to post-quantum cryptography has been an enormous challenge and effort for cryptographers over the last decade, with impressive results such as the future NIST standards. However, the latter has so far only considered central cryptographic mechanisms (signatures or KEM) and not more advanced ones, e.g., targeting privacy-preserving applications. Of particular interest is the family of solutions called blind signatures, group signatures and anonymous credentials, for which standards already exist, and which are deployed in billions of devices. Such a family does not have, at this stage, an efficient post-quantum counterpart although very recent works improved this state of affairs by offering two different alternatives: either one gets a system with rather large elements but a security proved under standard assumptions or one gets a more efficient system at the cost of ad-hoc interactive assumptions or weaker security models. Moreover, all these works have only considered size complexity without implementing the quite complex building blocks their systems are composed of. In other words, the practicality of such systems is still very hard to assess, which is a problem if one envisions a post-quantum transition for the corresponding systems/standards.
Sven Argo, Tim Güneysu, Corentin Jeudy, Georg Land, Adeline Roux-Langlois, Olivier Sanders
CCS2
2024 Three Sidekicks to Support Spectre Countermeasures
abstract
The Spectre attack revealed a critical security threat posed by speculative execution and since then numerous related attacks have been discovered and exploited to leak secrets across process boundaries. As the primary cause of the attack is deeply rooted in the microarchitectural processor design, mitigating speculative execution attacks with minimal impact on performance is far from straightforward. For example, various countermeasures have been proposed to limit speculative execution for certain instruction patterns, however, resulting in severe performance overheads. In this paper, we propose a set of code transformations to reduce the number of speculatively executed instructions and therefore significantly reduce the performance overhead of various countermeasures. We evaluate our code transformations combined with a hardware-based countermeasure in gem5. Our results demonstrate that our code transformations speed up the secure system by up to 16.6%.
Markus Krausz, Jan Philipp Thoma, Florian Stolz, Marc Fyrbiak, Tim Güneysu
DATE5
2024 Cips: The Cache Intrusion Prevention System
Jan Philipp Thoma, Florian Stolz, Tim Güneysu
ESORICS (4)3
2024 Agile Acceleration of Stateful Hash-based Signatures in Hardware
abstract
With the development of large-scale quantum computers, the current landscape of asymmetric cryptographic algorithms will change dramatically. Today’s standards like RSA, DSA, and ElGamal will no longer provide sufficient security against quantum attackers and need to be replaced with novel algorithms. In the face of these developments, NIST has already started a standardization process for new Key Encapsulation Mechanisms (KEMs) and Digital Signatures (DSs). Moreover, NIST has recommended the two stateful Hash-Based Signatures (HBSs) schemes XMSS and LMS for use in devices with a long expected lifetime and limited capabilities for maintenance. Both schemes are also standardized by the IETF. In this work, we present the first agile hardware implementation that supports both LMS and XMSS. Our design can instantiate either LMS, XMSS, or both schemes using a simple configuration setting. Leveraging the vast similarities of the two schemes, the hardware utilization of the agile design increases by 20% in LUTs and only 3% in Flip Flops (FFs) over a standalone XMSS implementation. Furthermore, our approach can easily be configured with an arbitrary number of hash cores and accelerators for the one-time signatures for different application scenarios. We evaluate our implementation on the Xilinx Artix-7 FPGA platform, which is the recommended target for PQC implementations by NIST. We explore potential tradeoffs in the design space and compare our results to previous work in this field.
Jan Philipp Thoma, Darius Hartlief, Tim Güneysu
ACM Trans. Embed. Comput. Syst.3
2023 Recommendation for a Holistic Secure Embedded ISA Extension
Florian Stolz, Marc Fyrbiak, Pascal Sasdrich, Tim Güneysu
ACNS4
2023 Implementing and Optimizing Matrix Triples with Homomorphic Encryption
abstract
In today’s interconnected world, data has become a valuable asset, leading to a growing interest in protecting it through techniques such as privacy-preserving computation. Two well-known approaches are multi-party computation and homomorphic encryption with use cases such as privacy-preserving machine learning evaluating or training neural networks. For multi-party computation, one of the fundamental arithmetic operations is the secure multiplication in the malicious security model and by extension the multiplication of matrices which is expensive to compute in the malicious model. Transferring the problem of secure matrix multiplication to the homomorphic domain enables savings in communication complexity, reducing the main bottleneck.
Johannes Mono, Tim Güneysu
AsiaCCS2
2023 Quantitative Fault Injection Analysis
Jakob Feldtkeller, Tim Güneysu, Patrick Schaumont
ASIACRYPT (4)2
2023 Combined Private Circuits - Combined Security Refurbished
abstract
Physical attacks are well-known threats to cryptographic implementations. While countermeasures against passive Side-Channel Analysis (SCA) and active Fault Injection Analysis (FIA) exist individually, protecting against their combination remains a significant challenge. A recent attempt at achieving joint security has been published at CCS 2022 under the name CINI-MINIS. The authors introduce relevant security notions and aim to construct arbitrary-order gadgets that remain trivially composable in the presence of a combined adversary. Yet, we show that all CINI-MINIS gadgets at any order are susceptible to a devastating attack with only a single fault and probe due to a lack of error correction modules in the compression. We explain the details of the attack, pinpoint the underlying problem in the constructions, propose an additional design principle, and provide new (fixed) provably secure and composable gadgets for arbitrary order. Luckily, the changes in the compression stage help us to save correction modules and registers elsewhere, making the resulting Combined Private Circuits (CPC) more secure and more efficient than the original ones. We also explain why the discovered flaws have been missed by the associated formal verification tool VERICA (TCHES 2022) and propose fixes to remove its blind spot. Finally, we explore alternative avenues to repair the compression stage without additional corrections based on non-completeness, i.e. constructing a compression that never recombines any secret. Yet, while this approach could have merit for low-order gadgets, it is, for now, hard to generalize and scales poorly to higher orders. We conclude that our refurbished arbitrary order CINI gadgets provide a solid foundation for further research.
Jakob Feldtkeller, Tim Güneysu, Thorben Moos, Jan Richter-Brockmann, Sayandeep Saha, Pascal Sasdrich, François-Xavier Standaert
CCS2
2023 EasiMask-Towards Efficient, Automated, and Secure Implementation of Masking in Hardware
abstract
Side-Channel Analysis (SCA) is a major threat to implementations of mathematically secure cryptographic algorithms. Applying masking countermeasures to hardware-based implementations is both time-consuming and error-prone due to side-effects buried deeply in the hardware design process. As a consequence, we propose our novel framework Easi-Mask in this work. Our semi-automated framework enables designers that have little experience with hardware implementation or physical security and the application of countermeasures to create a securely masked hardware implementation from an abstract description of a cryptographic algorithm. Its design-flow dismisses the developer from many challenges in the masking process of hardware implementations, while the generated implementations match the efficiency of hand-optimized designs from experienced security engineers. The modular approach can be mapped to arbitrary instantiations using different languages and transformations. We have verified the functionality, security, and efficiency of generated designs for several state of the art symmetric cryptographic algorithms, such as Advanced Encryption Standard (AES), Keccak, and PRESENT.
Fabian Buschkowski, Pascal Sasdrich, Tim Güneysu
DATE3
2023 PetaOps/W edge-AI $\mu$ Processors: Myth or reality?
abstract
With the rise of deep learning (DL), our world braces for artificial intelligence (AI) in every edge device, creating an urgent need for edge-AI SoCs. This SoC hardware needs to support high throughput, reliable and secure AI processing at ultra-low power (ULP), with a very short time to market. With its strong legacy in edge solutions and open processing platforms, the EU is well-positioned to become a leader in this SoC market. However, this requires AI edge processing to become at least 100 times more energy-efficient, while offering sufficient flexibility and scalability to deal with AI as a fast-moving target. Since the design space of these complex SoCs is huge, advanced tooling is needed to make their design tractable. The CONVOLVE project (currently in Inital stage) addresses these roadblocks. It takes a holistic approach with innovations at all levels of the design hierarchy. Starting with an overview of SOTA DL processing support and our project methodology, this paper presents 8 important design choices largely impacting the energy efficiency and flexibility of DL hardware. Finding good solutions is key to making smart-edge computing a reality.
Manil Dev Gomony, Floran de Putter, Anteneh Gebregiorgis, Gianna Paulin, Linyan Mei, Vikram Jain, Said Hamdioui, Victor Sanchez, Tobias Grosser, Marc Geilen, Marian Verhelst, Friedemann Zenke, Frank K. Gürkaynak, Barry de Bruin, Sander Stuijk, Simon Davidson, Sayandip De, Mounir Ghogho, Alexandra Jimborean, Sherif Eissa, Luca Benini, Dimitrios Soudris, Rajendra Bishnoi, Sam Ainsworth 0001, Federico Corradi, Ouassim Karrakchou, Tim Güneysu, Henk Corporaal
DATE27
2023 Dependability of Future Edge-AI Processors: Pandora's Box
abstract
This paper addresses one of the directions of the HORIZON EU CONVOLVE project being dependability of smart edge processors based on computation-in-memory and emerging memristor devices such as RRAM. It discusses how how this alternative computing paradigm will change the way we used to do manufacturing test. In addition, it describes how these emerging devices inherently suffering from many non-idealities are calling for new solutions in order to ensure accurate and reliable edge computing. Moreover, the paper also covers the security aspects for future edge processors and shows the challenges and the future directions.
Manil Dev Gomony, Anteneh Gebregiorgis, Moritz Fieback, Marc Geilen, Sander Stuijk, Jan Richter-Brockmann, Rajendra Bishnoi, Sven Argo, Lara Arche Andradas, Tim Güneysu, Mottaqiallah Taouil, Henk Corporaal, Said Hamdioui
ETS10
2023 Breaking and Protecting the Crystal: Side-Channel Analysis of Dilithium in Hardware
Hauke Malte Steffen, Georg Land, Lucie Johanna Kogelheide, Tim Güneysu
PQCrypto4
2023 SCARF - A Low-Latency Block Cipher for Secure Cache-Randomization
Federico Canale, Tim Güneysu, Gregor Leander, Jan Philipp Thoma, Yosuke Todo, Rei Ueno
USENIX Security Symposium2
2023 ClepsydraCache - Preventing Cache Attacks with Time-Based Evictions
Jan Philipp Thoma, Christian Niesler, Dominic A. Funke, Gregor Leander, Pierre Mayr, Nils Pohl, Lucas Davi, Tim Güneysu
USENIX Security Symposium8
2023 Revisiting Fault Adversary Models - Hardware Faults in Theory and Practice
abstract
Fault injection attacks are considered as powerful techniques to successfully attack embedded cryptographic implementations since various fault injection mechanisms from simple clock glitches to more advanced techniques like laser fault injection can lead to devastating attacks. Given these critical attack vectors, researchers came up with a long list of dedicated countermeasures to thwart such attacks. However, the security validation of proposed countermeasures is mostly performed on custom adversary models that are often not tightly coupled with the actual physical behavior of available fault injection mechanisms and, hence, fail to model the reality accurately. Furthermore, using custom models complicates comparison between different designs and evaluation results. As a consequence, we aim to close this gap by proposing a simple, generic, and consolidated fault injection adversary model that can be perfectly tailored to existing fault injection mechanisms and their physical behavior in hardware. To demonstrate the advantages, we apply it to a cryptographic primitive and evaluate it based on different attack vectors. We further show that our proposed adversary model can be integrated into the state-of-the-art fault verification tool VerFI. Finally, we provide a discussion on the benefits and differences of our approach compared to already existing evaluation methods.
Jan Richter-Brockmann, Pascal Sasdrich, Tim Güneysu
IEEE Trans. Computers3
2023 Challenges and Opportunities of Security-Aware EDA
abstract
The foundation of every digital system is based on hardware in which security, as a core service of many applications, should be deeply embedded. Unfortunately, the knowledge of system security and efficient hardware design is spread over different communities and, due to the complex and ever-evolving nature of hardware-based system security, state-of-the-art security is not always implemented in state-of-the-art hardware. However, automated security-aware hardware design seems to be a promising solution to bridge the gap between the different communities. In this work, we systematize state-of-the-art research with respect to security-aware Electronic Design Automation (EDA) and identify a modern security-aware EDA framework. As part of this work, we consider threats in the form of information flow, timing and power side channels, and fault injection, which are the fundamental building blocks of more complex hardware-based attacks. Based on the existing research, we provide important observations and research questions to guide future research in support of modern, holistic, and security-aware hardware design infrastructures.
Jakob Feldtkeller, Pascal Sasdrich, Tim Güneysu
ACM Trans. Embed. Comput. Syst.3
2022 Carry-Less to BIKE Faster
Ming-Shing Chen, Tim Güneysu, Markus Krausz, Jan Philipp Thoma
ACNS2
2022 CINI MINIS: Domain Isolation for Fault and Combined Security
abstract
Observation and manipulation of physical characteristics are well-known and powerful threats to cryptographic devices. While countermeasures against passive side-channel and active fault-injection attacks are well understood individually, combined attacks, i.e., the combination of fault injection and side-channel analysis, is a mostly unexplored area. Naturally, the complexity of analysis and secure construction increases with the sophistication of the adversary, making the combined scenario especially challenging. To tackle complexity, the side-channel community has converged on the construction of small building blocks, which maintain security properties even when composed. In this regard, Probe-Isolating Non-Interference (PINI) is a widely used notion for secure composition in the presence of side-channel attacks due to its efficiency and elegance. In this work, we transfer the core ideas behind PINI to the context of fault and combined security and, from that, construct the first trivially composable gadgets in the presence of a combined adversary.
Jakob Feldtkeller, Jan Richter-Brockmann, Pascal Sasdrich, Tim Güneysu
CCS4
2022 Proof-of-Possession for KEM Certificates using Verifiable Generation
abstract
Certificate authorities in public key infrastructures typically require entities to prove possession of the secret key corresponding to the public key they want certified. While this is straightforward for digital signature schemes, the most efficient solution for public key encryption and key encapsulation mechanisms (KEMs) requires an interactive challenge-response protocol, requiring a departure from current issuance processes. In this work we investigate how to non-interactively prove possession of a KEM secret key, specifically for lattice-based KEMs, motivated by the recently proposed KEMTLS protocol which replaces signature-based authentication in TLS 1.3 with KEM-based authentication. Although there are various zero-knowledge (ZK) techniques that can be used to prove possession of a lattice key, they yield large proofs or are inefficient to generate. We propose a technique called verifiable generation, in which a proof of possession is generated at the same time as the key itself is generated. Our technique is inspired by the Picnic signature scheme and uses the multi-party-computation-in-the-head (MPCitH) paradigm; this similarity to a signature scheme allows us to bind attribute data to the proof of possession, as required by certificate issuance protocols. We show how to instantiate this approach for two lattice-based KEMs in Round 3 of the NIST post-quantum cryptography standardization project, Kyber and FrodoKEM, and achieve reasonable proof sizes and performance. Our proofs of possession are faster and an order of magnitude smaller than the previous best MPCitH technique for knowledge of a lattice key, and in size-optimized cases can be comparable to even state-of-the-art direct lattice-based ZK proofs for Kyber. Our approach relies on a new result showing the uniqueness of Kyber and FrodoKEM secret keys, even if the requirement that all secret key components are small is partially relaxed, which may be of independent interest for improving efficiency of zero-knowledge proofs for other lattice-based statements.
Tim Güneysu, Philip W. Hodges, Georg Land, Mike Ounsworth, Douglas Stebila, Gregory M. Zaverucha
CCS1
2022 Efficiently Masking Polynomial Inversion at Arbitrary Order
Markus Krausz, Georg Land, Jan Richter-Brockmann, Tim Güneysu
PQCrypto4
2022 Write Me and I'll Tell You Secrets - Write-After-Write Effects On Intel CPUs
abstract
There is a long history of side channels in the memory hierarchy of modern CPUs. Especially the cache side channel is widely used in the context of transient execution attacks and covert channels. Therefore, many secure cache architectures have been proposed. Most of these architectures aim to make the construction of eviction sets infeasible by randomizing the address-to-cache mapping.
Jan Philipp Thoma, Tim Güneysu
RAID2
2022 Folding BIKE: Scalable Hardware Implementation for Reconfigurable Devices
abstract
Contemporary digital infrastructures and systems use and trust Public-Key Cryptography to exchange keys over insecure communication channels. With the development and progress in the research field of quantum computers, well established schemes like RSA and ECC are more and more threatened. The urgent demand to find and standardize new schemes – which are secure in a post-quantum world – was also realized by the National Institute of Standards and Technology which announced a Post-Quantum Cryptography Standardization Project in 2017. Recently, the round three candidates were announced and one of the alternate candidates is the Key Encapsulation Mechanism scheme BIKE. In this article, we investigate different strategies to efficiently implement the BIKE algorithm on Field-Programmable Gate Arrays (FPGAs). To this extend, we improve already existing polynomial multipliers, propose efficient strategies to realize polynomial inversions, and implement the Black-Gray-Flip decoder for the first time. Additionally, our implementation is designed to be scalable and generic with the BIKE specific parameters. All together, the fastest designs achieve latencies of 2.69 ms for the key generation, 0.1 ms for the encapsulation, and 1.89 ms for the decapsulation considering the lowest security level.
Jan Richter-Brockmann, Johannes Mono, Tim Güneysu
IEEE Trans. Computers3
2021 A Hard Crystal - Implementing Dilithium on Reconfigurable Hardware
Georg Land, Pascal Sasdrich, Tim Güneysu
CARDIS3
2021 Automated Masking of Software Implementations on Industrial Microcontrollers
abstract
Physical side-channel attacks threaten the security of exposed embedded devices, such as microcontrollers. Dedicated countermeasures, like masking, are necessary to prevent these powerful attacks. However, a gap between well-studied leakage models and observed leakage on real devices makes the application of these countermeasures non-trivial. This work provides a gadget-based concept to automated masking covering practically relevant leakage models to achieve security on real-world devices. We realize this concept with a fully automated compiler that transforms unprotected microcontroller-implementations of cryptographic primitives into masked executables, capable of being executed on the target device. In a case study, we apply our approach to a bitsliced LED implementation and perform a TVLA-based security evaluation of its core component: the PRESENT s-box.
Arnold Abromeit, Florian Bache, Leon A. Becker, Marc Gourjon, Tim Güneysu, Sabrina Jorn, Amir Moradi 0001, Maximilian Orlt, Falk Schellenberg
DATE5
2021 Nano Security: From Nano-Electronics to Secure Systems
abstract
The field of computer hardware stands at the verge of a revolution driven by recent breakthroughs in emerging nanodevices. “Nano Security” is a new Priority Program recently approved by DFG, the German Research Council. This initial-stage project initiative at the crossroads of nano-electronics and hardware-oriented security includes 11 projects with a total of 23 Principal Investigators from 18 German institutions. It considers the interplay between security and nano-electronics, focusing on a dichotomy which emerging nano-devices (and their architectural implications) have on system security. The projects within the Priority Program consider both: potential security threats and vulnerabilities stemming from novel nano-electronics, and innovative approaches to establishing and improving system security based on nano-electronics. This paper provides an overview of the Priority Program's overall philosophy and discusses the scientific objectives of its individual projects.
Ilia Polian, Frank Altmann, Tolga Arul, Christian Boit, Ralf Brederlow, Lucas Davi, Rolf Drechsler, Nan Du 0004, Thomas Eisenbarth 0001, Tim Güneysu, Sascha Hermann, Matthias Hiller, Rainer Leupers, Farhad Merchant, Thomas Mussenbrock, Stefan Katzenbeisser 0001, Akash Kumar 0001, Wolfgang Kunz, Thomas Mikolajick, Vivek Pachauri, Jean-Pierre Seifert, Frank Sill, Jens Trommer
DATE10
2021 BasicBlocker: ISA Redesign to Make Spectre-Immune CPUs Faster
abstract
Recent research has revealed an ever-growing class of microarchitectural attacks that exploit speculative execution, a standard feature in modern processors. Proposed and deployed countermeasures involve a variety of compiler updates, firmware updates, and hardware updates. None of the deployed countermeasures have convincing security arguments, and many of them have already been broken.
Jan Philipp Thoma, Jakob Feldtkeller, Markus Krausz, Tim Güneysu, Daniel J. Bernstein
RAID4
2020 Concurrent error detection revisited: hardware protection against fault and side-channel attacks
abstract
Fault Injection Analysis (FIA) and Side-Channel Analysis (SCA) are considered among the most serious threats to cryptographic implementations and require dedicated countermeasures to ensure protection through the entire life-cycle of the implementations.
Jan Richter-Brockmann, Pascal Sasdrich, Florian Bache, Tim Güneysu
ARES4
2020 Improved Side-Channel Resistance by Dynamic Fault-Injection Countermeasures
abstract
Side-channel analysis and fault-injection attacks are known as serious threats to cryptographic hardware implementations and the combined protection against both is currently an open line of research. A promising countermeasure with considerable implementation overhead appears to be a mix of first-order secure Threshold Implementations and linear Error-Correcting Codes.In this paper we employ for the first time the inherent structure of non-systematic codes as fault countermeasure which dynamically mutates the applied generator matrices to achieve a higher-order side-channel and fault-protected design. As a case study, we apply our scheme to the PRESENT block cipher that do not show any higher-order side-channel leakage after measuring 150 million power traces.
Jan Richter-Brockmann, Tim Güneysu
ASAP2
2020 Revisiting ECM on GPUs
Jonas Wloka, Jan Richter-Brockmann, Colin Stahlke, Thorsten Kleinjung, Christine Priplata, Tim Güneysu
CANS6
2020 Deep Learning Multi-Channel Fusion Attack Against Side-Channel Protected Hardware
abstract
State-of-the-art hardware masking approaches like threshold implementations and domain-oriented masking provide a guaranteed level of security even in the presence of glitches. Although provable secure in theory, recent work showed that the effective security order of a masked hardware implementation can be lowered by applying a multi-probe attack or exploiting externally amplified coupling effects. However, the proposed attacks are based on an unrealistic adversary model (i.e. knowledge of masks values during profiling) or require complex measurement setup manipulations.In this work, we propose a novel attack vector that exploits location dependent leakage from several decoupling capacitors of a modern System-on-Chip (SoC) with 16 nm fabrication technology. We combine the leakage from different sources using a deep learning-based information fusion approach. The results show a remarkable advantage regarding the number of required traces for a successful key recovery compared to state-of-the-art profiled side-channel attacks. All evaluations are performed under realistic conditions, resulting in a real-world attack scenario that is not limited to academic environments.
Benjamin Hettwer, Daniel Fennes, Sebastien Leger, Jan Richter-Brockmann, Stefan Gehrer, Tim Güneysu
DAC6
2020 Towards Secure Composition of Integrated Circuits and Electronic Systems: On the Role of EDA
abstract
Modern electronic systems become evermore complex, yet remain modular, with integrated circuits (ICs) acting as versatile hardware components at their heart. Electronic design automation (EDA) for ICs has focused traditionally on power, performance, and area. However, given the rise of hardware-centric security threats, we believe that EDA must also adopt related notions like secure by design and secure composition of hardware. Despite various promising studies, we argue that some aspects still require more efforts, for example: effective means for compilation of assumptions and constraints for security schemes, all the way from the system level down to the "bare metal"; modeling, evaluation, and consideration of security-relevant metrics; or automated and holistic synthesis of various countermeasures, without inducing negative cross-effects.In this paper, we first introduce hardware security for the EDA community. Next we review prior (academic) art for EDA-driven security evaluation and implementation of countermeasures. We then discuss strategies and challenges for advancing research and development toward secure composition of circuits and systems.
Johann Knechtel, Elif Bilge Kavun, Francesco Regazzoni 0001, Annelie Heuser, Anupam Chattopadhyay, Debdeep Mukhopadhyay, Soumyajit Dey, Yunsi Fei, Yaacov Belenky, Itamar Levi, Tim Güneysu, Patrick Schaumont, Ilia Polian
DATE11
2020 Lightweight Side-Channel Protection using Dynamic Clock Randomization
abstract
Power analysis attacks have evolved rapidly over the past two decades, recently strengthened by advanced deep learning algorithms. However, the application of effective countermeasures such as masking is often challenging in practice due to restricted power and area resources of cryptographic devices. On the other hand, lightweight hiding methods like random data delays often introduce only a small amount of entropy in the execution process, and thus provide only a moderate level of protection. In this work, we propose and evaluate a generic hiding countermeasure based on dynamic clock frequency randomization. We exploit runtime reconfiguration of modern reconfigurable devices to produce a highly unstable clock signal, which yields up to 20 million different execution times for an AES encryption operation. Our design not only creates heavy misalignments in the power traces, but is also highly customizable and can be easily composed with other side-channel countermeasures. We test our approach using recently proposed evaluation methods for desynchronized power traces including sliding-window correlation analysis and deep neural networks. The results show that none of the attacks is able to recover the secret key with one million power traces. Furthermore, we could not detect any first-order leakage in five million encryptions using state-of-the-art leakage assessment.
Benjamin Hettwer, Kallyan Das, Sebastien Leger, Stefan Gehrer, Tim Güneysu
FPL5
2020 Verification of Embedded Binaries using Coverage-guided Fuzzing with SystemC-based Virtual Prototypes
abstract
Extensive verification of embedded SW is very important to avoid errors and security vulnerabilities. Therefore, mainly simulation-based methods are employed that leverage Virtual Prototypes (VPs) for SW execution early in the design flow. VPs are essentially abstract models of the entire HW platform including peripherals. They are predominantly created in SystemC. However, a comprehensive simulation-based verification requires integration of sophisticated test generation techniques.
Vladimir Herdt, Daniel Große, Jonas Wloka, Tim Güneysu, Rolf Drechsler
ACM Great Lakes Symposium on VLSI4
2019 Securing Cryptographic Circuits by Exploiting Implementation Diversity and Partial Reconfiguration on FPGAs
abstract
Adaptive and reconfigurable systems such as Field Programmable Gate Arrays (FPGAs) play an integral part of many complex embedded platforms. This implies the capability to perform runtime changes to hardware circuits on demand. In this work, we make use of this feature to propose a no-vel countermeasure against physical attacks of cryptographic implementations. In particular, we leverage exploration of the implementation space on FPGAs to create various circuits with different hardware layouts from a single design of the Advanced Encryption Standard (AES), that are dynamically exchanged during device operation. We provide evidence from practical experiments based on a modern Xilinx ZYNQ UltraScale+ FPGA that our approach increases the resistance against physical attacks by at least factor two. Furthermore, the genericness of our approach allows an easy adaption to other algorithms and combination with other countermeasures.
Benjamin Hettwer, Johannes Petersen, Stefan Gehrer, Heike Neumann, Tim Güneysu
DATE5
2019 Towards Practical Microcontroller Implementation of the Signature Scheme Falcon
Tobias Oder, Julian Speith, Kira Höltgen, Tim Güneysu
PQCrypto4
2019 Deep Neural Network Attribution Methods for Leakage Analysis and Symmetric Key Recovery
Benjamin Hettwer, Stefan Gehrer, Tim Güneysu
SAC3
2018 Confident leakage assessment - A side-channel evaluation framework based on confidence intervals
abstract
Cryptographic devices that potentially operate in hostile physical environments need to be secured against side-channel attacks. In order to ensure the effectiveness of the required countermeasures, scientists, developers, and evaluators need efficient methods to test the security level of a device. In this paper we propose a new framework based on confidence intervals that extends established t-test based approaches for test-vector leakage assessment (TVLA). In comparison to previous TVLA approaches the new methodology does not only enable the detection of leakage but can also assert its absence. The framework is robust against noise in the evaluation system and thereby avoids false negatives. These improvements can be achieved without overhead in measurement complexity and with a minimum of additional computational costs compared to previous approaches. We evaluate our method under realistic conditions by applying it to a protected implementation of AES.
Florian Bache, Christina Plump, Tim Güneysu
DATE3
2018 Physical Protection of Lattice-Based Cryptography: Challenges and Solutions
abstract
The impending realization of scalable quantum computers will have a significant impact on today's security infrastructure. With the advent of powerful quantum computers public key cryptographic schemes will become vulnerable to Shor's quantum algorithm, undermining the security current communications systems. Post-quantum (or quantum-resistant) cryptography is an active research area, endeavoring to develop novel and quantum resistant public key cryptography. Amongst the various classes of quantum-resistant cryptography schemes, lattice-based cryptography is emerging as one of the most viable options. Its efficient implementation on software and on commodity hardware has already been shown to compete and even excel the performance of current classical security public-key schemes. This work discusses the next step in terms of their practical deployment, i.e., addressing the physical security of lattice-based cryptographic implementations. We survey the state-of-the-art in terms of side channel attacks (SCA), both invasive and passive attacks, and proposed countermeasures. Although the weaknesses exposed have led to countermeasures for these schemes, the cost, practicality and effectiveness of these on multiple implementation platforms, however, remains under-studied.
Ayesha Khalid, Tobias Oder, Felipe Valencia, Máire O'Neill, Tim Güneysu, Francesco Regazzoni 0001
ACM Great Lakes Symposium on VLSI5
2018 Profiled Power Analysis Attacks Using Convolutional Neural Networks with Domain Knowledge
Benjamin Hettwer, Stefan Gehrer, Tim Güneysu
SAC3
2018 GliFreD: Glitch-Free Duplication Towards Power-Equalized Circuits on FPGAs
abstract
Designers of secure hardware are required to harden their implementations against physical threats, such as power analysis attacks. In particular, cryptographic hardware circuits need to decorrelate their current consumption from the information inferred by processing (secret) data. A common technique to achieve this goal is the use of special logic styles that aim at equalizing the current consumption at each single processing step. However, since all hiding techniques like Dual-Rail Precharge (DRP) were originally developed for ASICs, the deployment of such countermeasures on FPGA devices with fixed and predefined logic structure poses a particular challenge. In this work, we propose and practically evaluate a new DRP scheme (GliFreD) that has been exclusively designed for FPGA platforms. GliFreD overcomes the well-known early propagation issue, prevents glitches, uses an isolated dual-rail concept, and mitigates imbalanced routings. With all these features, GliFreD significantly exceeds the level of physical security achieved by any previously reported, related countermeasures for FPGAs.
Alexander Wild, Amir Moradi 0001, Tim Güneysu
IEEE Trans. Computers3
2017 Hiding Higher-Order Side-Channel Leakage - Randomizing Cryptographic Implementations in Reconfigurable Hardware
Pascal Sasdrich, Amir Moradi 0001, Tim Güneysu
CT-RSA3
2017 Cryptography for Next Generation TLS: Implementing the RFC 7748 Elliptic Curve448 Cryptosystem in Hardware
abstract
With RFC 7748 the two elliptic curves Curve25519 and Curve448 were proposed for the next generation of TLS. Both curves were designed and optimized purely for software implementation; their implementation in hardware or physical protection against side-channel attacks were not considered in the design phase. Recently, it has been shown that for Curve25519 an efficient implementations in hardware along with side-channel protection is feasible -- yet results for the high-security Curve448 are missing. In this work we demonstrate that Curve448 can indeed be efficiently and securely implemented in hardware. We present a novel architecture for Curve448 that can compute more than 1000 point multiplications per second with 1580 logic slices and 33 DSP units of a Xilinx XC7Z020 FPGA.
Pascal Sasdrich, Tim Güneysu
DAC2
2017 SPARX - A side-channel protected processor for ARX-based cryptography
abstract
ARX-based cryptographic algorithms are composed of only three elemental operations — addition, rotation and exclusive or — which are mixed to ensure adequate confusion and diffusion properties. While ARX-ciphers can easily be protected against timing attacks, special measures like masking have to be taken in order to prevent power and electromagnetic analysis. In this paper we present a processor architecture for ARX-based cryptography, that intrinsically guarantees first-order SCA resistance of any implemented algorithm. This is achieved by protecting the complete data path using a Boolean masking scheme with three shares. We evaluate our security claims by mapping an ARX-algorithm to the proposed architecture and using the common leakage detection methodology based on Student's i-test to certify the side-channel resistance of our processor.
Florian Bache, Tobias Schneider 0002, Amir Moradi 0001, Tim Güneysu
DATE4
2017 A fair and comprehensive large-scale analysis of oscillation-based PUFs for FPGAs
abstract
Physical Unclonable Functions (PUFs) have gained a lot of research attention in recent years resulting in many different PUF proposals. Several of these proposals were aimed specifically at FPGA implementations. However, often these PUFs are evaluated and implemented for different (and often old) FPGA families with different metrics. Missing implementation details in many papers further hamper a fair analysis, as small details such as the exact routing can have significant impact on the PUF performance. In this paper we aim to overcome these problems by providing a fair comparison of some of the most promising Weak PUFs for FPGAs, the classic Ring Oscillator PUF (RO PUF), the Loop PUF and the TERO PUF. Each PUF is implemented with the same area optimizations and careful manual routing for modern Xilinx Artix-7 FPGAs and several implementation options are discussed. We measure the reliability and uniqueness of the PUF constructs on 100 BASYS-3 boards for a temperature range of -22°C to 44°C and use a glitch-generating core to analyze the vulnerability of the PUF constructs to surrounding logic. Our results show that the RO PUF has the best reliability in the presence of temperature variations while TERO has the best uniqueness of the three considered PUFs. Interestingly, the TERO PUF also shows the highest resistance to surrounding logic. To encourage further research in FPGA PUFs and to enable a fair comparison to future work the implementations as well as the measurement data will be made publicly available.
Alexander Wild, Georg T. Becker, Tim Güneysu
FPL3
2017 CAKE: Code-Based Algorithm for Key Encapsulation
Paulo S. L. M. Barreto, Shay Gueron, Tim Güneysu, Rafael Misoczki, Edoardo Persichetti, Nicolas Sendrier, Jean-Pierre Tillich
IMACC3
2017 High-Performance Ideal Lattice-Based Cryptography on 8-Bit AVR Microcontrollers
abstract
Over recent years lattice-based cryptography has received much attention due to versatile average-case problems like Ring-LWE or Ring-SIS that appear to be intractable by quantum computers. In this work, we evaluate and compare implementations of Ring-LWE encryption and the bimodal lattice signature scheme (BLISS) on an 8-bit Atmel ATxmega128 microcontroller. Our implementation of Ring-LWE encryption provides comprehensive protection against timing side-channels and takes 24.9ms for encryption and 6.7ms for decryption. To compute a BLISS signature, our software takes 317ms and 86ms for verification. These results underline the feasibility of lattice-based cryptography on constrained devices.
Zhe Liu 0001, Thomas Pöppelmann, Tobias Oder, Hwajeong Seo, Sujoy Sinha Roy, Tim Güneysu, Johann Großschädl, Howon Kim 0001, Ingrid Verbauwhede
ACM Trans. Embed. Comput. Syst.6
2016 A grain in the silicon: SCA-protected AES in less than 30 slices
abstract
AES is the predominant block cipher used worldwide in many cryptographic applications. Despite of the wealth of already available implementations, we here introduce an ultra-lightweight AES-128 implementation specifically tailored for reconfigurable hardware. Our basic proposal presents a full AES-128 providing 9.12 Mbit/s throughput and occupying just 21 slices of a Spartan-6 and no additional memories. We also show that this architecture almost, inherently supports shuffling as side-channel countermeasure and provide results of a practical evaluation. Our protected design fits into 24 slices providing 7.82 Mbit/s throughput. Finally, we present a complete AES core that combines previous results with random number generation which fits 28 slices at 4.35 Mbit/s throughput.
Pascal Sasdrich, Tim Güneysu
ASAP2
2016 Sixth International Workshop on Trustworthy Embedded Devices (TrustED 2016)
abstract
The Internet of Things (IoT) is expected to become a global information and communication infrastructure for cyber physical systems and to bring numerous value-added services for modern society. However, the integration of heterogeneous devices and service models into a cohesive system significantly increases the complexity of design and deployment and introduces the new challenges for the security of systems and processed as well as the privacy of the collected data. The Workshop on Trustworthy Embedded Devices (TrustED) addresses all aspects of security and privacy related to embedded systems and the IoT. TrustED 2016 is a continuation of previous workshops in this series, which were held in conjunction with ESORICS 2011, IEEE Security & Privacy 2012, ACM CCS 2013, ACM CCS 2014, and ACM CCS 2015 (see http://www.trusted-workshop.de for details). The goal of this workshop is to bring together experts from academia and research institutes, industry, and government in the field of security and privacy in cyber physical systems to discuss and investigate the problems, challenges, and recent scientific and technological developments.
Xinxin Fan, Tim Güneysu
CCS2
2016 Strong 8-bit Sboxes with Efficient Masking in Hardware
Erik Boss, Vincent Grosso, Tim Güneysu, Gregor Leander, Amir Moradi 0001, Tobias Schneider 0002
CHES3
2016 ParTI - Towards Combined Hardware Countermeasures Against Side-Channel and Fault-Injection Attacks
Tobias Schneider 0002, Amir Moradi 0001, Tim Güneysu
CRYPTO (2)3
2016 Standard lattices in hardware
abstract
Lattice-based cryptography has gained credence recently as a replacement for current public-key cryptosystems, due to its quantum-resilience, versatility, and relatively low key sizes. To date, encryption based on the learning with errors (LWE) problem has only been investigated from an ideal lattice standpoint, due to its computation and size efficiencies. However, a thorough investigation of standard lattices in practice has yet to be considered. Standard lattices may be preferred to ideal lattices due to their stronger security assumptions and less restrictive parameter selection process.
James Howe, Ciara Rafferty, Máire O'Neill, Francesco Regazzoni 0001, Tim Güneysu, K. Beeden
DAC5
2016 White-Box Cryptography in the Gray Box - - A Hardware Implementation and its Side Channels -
Pascal Sasdrich, Amir Moradi 0001, Tim Güneysu
FSE3
2016 IND-CCA Secure Hybrid Encryption from QC-MDPC Niederreiter
Ingo von Maurich, Lukas Heberle, Tim Güneysu
PQCrypto3
2016 Bridging the Gap: Advanced Tools for Side-Channel Leakage Estimation Beyond Gaussian Templates and Histograms
Tobias Schneider 0002, Amir Moradi 0001, François-Xavier Standaert, Tim Güneysu
SAC4
2016 Information reconciliation schemes in physical-layer security: A survey
Christopher Huth, René Guillaume, Thomas Strohm, Paul Duplys, Irin Ann Samuel, Tim Güneysu
Comput. Networks6
2015 Arithmetic Addition over Boolean Masking - Towards First- and Second-Order Resistance in Hardware
Tobias Schneider 0002, Amir Moradi 0001, Tim Güneysu
ACNS3
2015 New ASIC/FPGA Cost Estimates for SHA-1 Collisions
abstract
SHA-1 remains, till date, the most widely used hash function, in spite of several successful cryptanalytic attacks against it. These attacks, however, remain impractical due to high computation complexity and associated cost. We endeavor to do cost-time product estimation for an attack by the aid of application-specific hardware acceleration. This work proposes an Application-Specific Instruction-set Processor (ASIP), named Cracken. Cracken is aimed to efficiently realize near collision attack on SHA-1. The estimations of the physical attack complexity is done using 65nm standard CMOS technology and commercial FPGA devices. It is estimated, with post-layout simulations, that Stevens' differential attack with an estimated complexity of 2^57.5, can be executed in 46 days using 4096 Cracken cores at a cost of Euros 15m. Estimation for real collision with complexity 2^61 is also done. Our cost-time estimates reveal that an FPGA-based attack is more efficient compared to ASIC. Previously reported SHA-1 attacks based on ASIC and cloud computing platforms are also compiled and benchmarked for reference.
Ayesha Khalid, Anupam Chattopadhyay, Christian Rechberger, Tim Güneysu, Christof Paar
DSD5
2015 Affine Equivalence and Its Application to Tightening Threshold Implementations
Pascal Sasdrich, Amir Moradi 0001, Tim Güneysu
SAC3
2015 Lattice-Based Signatures: Optimization and Implementation on Reconfigurable Hardware
abstract
Nearly all of the currently used signature schemes, such as RSA or DSA, are based either on the factoring assumption or the presumed intractability of the discrete logarithm problem. As a consequence, the appearance of quantum computers or algorithmic advances on these problems may lead to the unpleasant situation that a large number of today's schemes will most likely need to be replaced with more secure alternatives. In this work we present such an alternative-an efficient signature scheme whose security is derived from the hardness of lattice problems. It is based on recent theoretical advances in lattice-based cryptography and is highly optimized for practicability and use in embedded systems. The public and secret keys are roughly 1.5 kB and 0.3 kB long, while the signature size is approximately 1.1 kB for a security level of around 80 bits. We provide implementation results on reconfigurable hardware (Spartan/Virtex-6) and demonstrate that the scheme is scalable, has low area consumption, and even outperforms classical schemes.
Tim Güneysu, Vadim Lyubashevsky, Thomas Pöppelmann
IEEE Trans. Computers1
2015 Practical Lattice-Based Digital Signature Schemes
abstract
Digital signatures are an important primitive for building secure systems and are used in most real-world security protocols. However, almost all popular signature schemes are either based on the factoring assumption (RSA) or the hardness of the discrete logarithm problem (DSA/ECDSA). In the case of classical cryptanalytic advances or progress on the development of quantum computers, the hardness of these closely related problems might be seriously weakened. A potential alternative approach is the construction of signature schemes based on the hardness of certain lattice problems that are assumed to be intractable by quantum computers. Due to significant research advancements in recent years, lattice-based schemes have now become practical and appear to be a very viable alternative to number-theoretic cryptography. In this article, we focus on recent developments and the current state of the art in lattice-based digital signatures and provide a comprehensive survey discussing signature schemes with respect to practicality. Additionally, we discuss future research areas that are essential for the continued development of lattice-based cryptography.
James Howe, Thomas Pöppelmann, Máire O'Neill, Elizabeth O'Sullivan, Tim Güneysu
ACM Trans. Embed. Comput. Syst.5
2015 Implementing QC-MDPC McEliece Encryption
abstract
With respect to performance, asymmetric code-based cryptography based on binary Goppa codes has been reported as a highly interesting alternative to RSA and ECC. A major drawback is still the large keys in the range between 50 and 100KB that prevented real-world applications of code-based cryptosystems so far. A recent proposal by Misoczki et al. showed that quasi-cyclic moderate-density parity-check (QC-MDPC) codes can be used in McEliece encryption, reducing the public key to just 0.6KB to achieve an 80-bit security level. In this article, we provide optimized decoding techniques for MDPC codes and survey several efficient implementations of the QC-MDPC McEliece cryptosystem. This includes high-speed and lightweight architectures for reconfigurable hardware, efficient coding styles for ARM’s Cortex-M4 microcontroller, and novel high-performance software implementations that fully employ vector instructions. Finally, we conclude that McEliece encryption in combination with QC-MDPC codes not only enables high-performance implementations but also allows for lightweight designs on a wide range of different platforms.
Ingo von Maurich, Tobias Oder, Tim Güneysu
ACM Trans. Embed. Comput. Syst.3
2015 Introduction for Embedded Platforms for Cryptography in the Coming Decade
abstract
editorial Free Access Share on Introduction for Embedded Platforms for Cryptography in the Coming Decade Editors: Patrick Schaumont Virginia Tech, USA Virginia Tech, USAView Profile , Maire O'Neill Queen's University Belfast, United Kingdom Queen's University Belfast, United KingdomView Profile , Tim Güneysu Ruhr University Bochum, Germany Ruhr University Bochum, GermanyView Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 14Issue 3May 2015 Article No.: 40pp 1–3https://doi.org/10.1145/2745710Published:21 April 2015Publication History 2citation284DownloadsMetricsTotal Citations2Total Downloads284Last 12 Months15Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Patrick Schaumont, Máire O'Neill, Tim Güneysu
ACM Trans. Embed. Comput. Syst.3
2015 Implementing Curve25519 for Side-Channel-Protected Elliptic Curve Cryptography
abstract
For security-critical embedded applications Elliptic Curve Cryptography (ECC) has become the predominant cryptographic system for efficient key agreement and digital signatures. However, ECC still involves complex modular arithmetic that is a particular burden for small processors. In this context, Bernstein proposed the highly efficient ECC instance Curve25519 that particularly enables efficient software implementations at a security level comparable to AES-128 with inherent resistance to simple power analysis (SPA) and timing attacks. In this work, we show that Curve25519 is likewise competitive on FPGAs even when countermeasures to thwart side-channel power analysis are included. Our basic multicore DSP-based architectures achieves a maximal performance of more than 32,000 point multiplications per second on a Xilinx Zynq 7020 FPGA. Including a mix of side-channel countermeasures to impede simple and differential power analysis, we still achieve more than 27,500 point multiplications per second with a moderate increase in logic resources.
Pascal Sasdrich, Tim Güneysu
ACM Trans. Reconfigurable Technol. Syst.2
2014 Enhanced Lattice-Based Signatures on Reconfigurable Hardware
Thomas Pöppelmann, Léo Ducas, Tim Güneysu
CHES3
2014 Beyond ECDSA and RSA: Lattice-based Digital Signatures on Constrained Devices
abstract
All currently deployed asymmetric cryptography is broken with the advent of powerful quantum computers. We thus have to consider alternative solutions for systems with long-term security requirements (e.g., for long-lasting vehicular and avionic communication infrastructures). In this work we present an efficient implementation of BLISS, a recently proposed, post-quantum secure, and formally analyzed novel lattice-based signature scheme. We show that we can achieve a significant performance of 35.3 and 6 ms for signing and verification, respectively, at a 128-bit security level on an ARM Cortex-M4F microcontroller. This shows that lattice-based cryptography can be efficiently deployed on today's hardware and provides security solutions for many use cases that can even withstand future threats.
Tobias Oder, Thomas Pöppelmann, Tim Güneysu
DAC3
2014 Lightweight code-based cryptography: QC-MDPC McEliece encryption on reconfigurable devices
abstract
With the break of RSA and ECC cryptosystems in an era of quantum computing, asymmetric code-based cryptography is an established alternative that can be a potential replacement. A major drawback are large keys in the range between 50kByte to several MByte that prevented real-world applications of code-based cryptosystems so far. A recent proposal by Misoczki et al. showed that quasi-cyclic moderate density parity-check (QC-MDPC) codes can be used in McEliece encryption — reducing the public key to just 0.6 kByte to achieve a 80-bit security level. Despite of reasonably small key sizes that could also enable small designs, previous work only report highperformance implementations with high resource consumptions of more than 13,000 slices on a large Xilinx Virtex-6 FPGA for a combined en-/decryption unit. In this work we focus on lightweight implementations of code-based cryptography and demonstrate that McEliece encryption using QC-MDPC codes can be implemented with a significantly smaller resource footprint — still achieving reasonable performance sufficient for many applications, e.g., challenge-response protocols or hybrid firmware encryption. More precisely, our design requires just 68 slices for the encryption and around 150 slices for the decryption unit and is able to en-/decrypt an input block in 2.2ms and 13.4 ms, respectively.
Ingo von Maurich, Tim Güneysu
DATE2
2014 Fault Sensitivity Analysis Meets Zero-Value Attack
abstract
Previous works have shown that the combinatorial path delay of a cryptographic function, e.g., The AES S-box, depends on its input value. Since the relation between critical path delay and input value seems to be relatively random and highly dependent on the routing of the circuit, up to now only template or some collision attacks could reliably extract the used secret key of implementations not protected against fault attacks. Here we present a new attack which is based on the fact that, because of the zero-to-zero mapping of the AES Sbox inversion circuit, the critical path when processing the zero input is notably shorter than for all other inputs. Applying the attack to an AES design protected by an state-of-the-art fault detection scheme, we are able to fully recover the secret key in less than eight hours. Note that we neither require a known key measurement step (template case) nor a high similarity between different S-box instances (collision case). The only information gathered from the device is whether a fault occurred when processing a chosen plaintext.
Oliver Mischke, Amir Moradi 0001, Tim Güneysu
FDTC3
2014 THOR - The hardware onion router
abstract
Security and privacy of data traversing internet have always been a major concern for all users. In this context, The Onion Routing (Tor) is the most successful protocol to anonymize global Internet traffic and is widely deployed as software on many personal computers or servers. In this paper, we explore the potential of modern reconfigurable devices to efficiently realize the Tor protocol on embedded devices. In particular, this targets the acceleration of the complex cryptographic operations involved in the handshake of routing nodes and the data stream encryption. Our hardware-based implementation on the Xilinx Zynq platform outperforms previous embedded solutions by more than a factor of 9 with respect to the cryptographic handshake - ultimately enabling quite inexpensive but highly efficient routers. Hence, we consider our work as a further milestone towards the development and the dissemination of low-cost and high performance onion relays that hopefully ultimately leads again to a more private Internet.
Tim Güneysu, Francesco Regazzoni 0001, Pascal Sasdrich, Marcin Wójcik
FPL1
2014 Enabling SRAM-PUFs on Xilinx FPGAs
abstract
Physically Unclonable Functions (PUFs) based on the evaluation of uninitialized SRAM are one of the most promising PUF candidates to date. However, transferring their concept to Xilinx FPGAs is not straightforward since all SRAM-based block memories in these FPGAs are automatically cleared on power-up, destroying the desired initial bits of information. In this work we therefore propose a novel strategy to convert block memories of 28nm Xilinx FPGAs into SRAM-PUFs by exploiting their recently introduced feature of power-gating and partial reconfiguration.
Alexander Wild, Tim Güneysu
FPL2
2014 Area optimization of lightweight lattice-based encryption on reconfigurable hardware
abstract
Ideal lattice-based cryptography gained significant attraction in the last years due to its versatility, simplicity and performance in implementations. Nevertheless, existing implementations of encryption schemes reported only results trimmed for high-performance what is certainly not sufficient for all applications in practice. To the contrary, in this work we investigate lightweight aspects and suitable parameter sets for Ring-LWE encryption and show optimizations that enable implementations even with very few resources on a reconfigurable hardware device. Despite of this restriction, we still achieve reasonable throughput that is sufficient for many today's and future applications.
Thomas Pöppelmann, Tim Güneysu
ISCAS2
2014 Towards Side-Channel Resistant Implementations of QC-MDPC McEliece Encryption on Constrained Devices
Ingo von Maurich, Tim Güneysu
PQCrypto2
2013 Attacking Atmel's CryptoMemory EEPROM with Special-Purpose Hardware
Alexander Wild, Tim Güneysu, Amir Moradi 0001
ACNS2
2013 Efficient implementation of cryptographic primitives on the GA144 multi-core architecture
abstract
With myriads of small and pervasive devices in our digital age, the availability of low-power and energy-efficient processing technology has become absolutely essential. Most of these constrained devices need to incorporate security services for confidentiality and privacy in addition to their primary tasks - typically involving computationally expensive cryptography. In the last years, many researchers have worked on novel lightweight cryptographic constructions to minimize the computational burden on the constrained devices. However, most of those alternative constructions sacrificed security for simplicity, potentially enabling just as simple attacks. In this work, we aim for another approach and implement standardized and well-established cryptography on a special but very lightweight platform, namely an asynchronous GA144 ultra-low-powered multicore processor with 144 simplistic cores. For the first time, we demonstrate that symmetric and asymmetric cryptography such as AES and RSA is even feasible on such a low-end and unclocked device. With energy consumption being as low as 0.63 μJ and 22.3 mJ, this platform achieves a performance of 38 μs and 462.9 ms per AES and RSA operation, respectively. Both energy consumption as well as computation time are significantly lower than many lightweight implementations reported so far.
Tobias Schneider 0002, Ingo von Maurich, Tim Güneysu
ASAP3
2013 Smaller Keys for Code-Based Cryptography: QC-MDPC McEliece Implementations on Embedded Devices
Stefan Heyse, Ingo von Maurich, Tim Güneysu
CHES3
2013 Software Speed Records for Lattice-Based Signatures
Tim Güneysu, Tobias Oder, Thomas Pöppelmann, Peter Schwabe
PQCrypto1
2013 Towards Practical Lattice-Based Public-Key Encryption on Reconfigurable Hardware
Thomas Pöppelmann, Tim Güneysu
Selected Areas in Cryptography2
2012 PRINCE - A Low-Latency Block Cipher for Pervasive Computing Applications - Extended Abstract
Julia Borghoff, Anne Canteaut, Tim Güneysu, Elif Bilge Kavun, Miroslav Knezevic, Lars R. Knudsen, Gregor Leander, Ventzislav Nikov, Christof Paar, Christian Rechberger, Peter Rombouts, Søren S. Thomsen, Tolga Yalçin
ASIACRYPT3
2012 Compact Implementation and Performance Evaluation of Hash Functions in ATtiny Devices
Josep Balasch, Baris Ege, Thomas Eisenbarth 0001, Benoît Gérard, Tim Güneysu, Stefan Heyse, Stéphanie Kerckhof, François Koeune, Thomas Plos, Thomas Pöppelmann, Francesco Regazzoni 0001, François-Xavier Standaert, Gilles Van Assche, Ronny Van Keer, Loïc van Oldeneel tot Oldenzeel, Ingo von Maurich
CARDIS6
2012 Practical Lattice-Based Cryptography: A Signature Scheme for Embedded Systems
Tim Güneysu, Vadim Lyubashevsky, Thomas Pöppelmann
CHES1
2012 Towards One Cycle per Bit Asymmetric Encryption: Code-Based Cryptography on Reconfigurable Hardware
Stefan Heyse, Tim Güneysu
CHES2
2012 Evaluation of Standardized Password-Based Key Derivation against Parallel Processing Platforms
Markus Dürmuth, Tim Güneysu, Markus Kasper, Christof Paar, Tolga Yalçin, Ralf Zimmermann 0001
ESORICS2
2011 An Experimentally Verified Attack on Full Grain-128 Using Dedicated Reconfigurable Hardware
Itai Dinur, Tim Güneysu, Christof Paar, Adi Shamir, Ralf Zimmermann 0001
ASIACRYPT2
2011 Generic Side-Channel Countermeasures for Reconfigurable Devices
Tim Güneysu, Amir Moradi 0001
CHES1
2011 The future of high-speed cryptography: new computing platforms and new ciphers
abstract
Almost all of today's security systems rely on cryptographic primitives as core components which are usually considered the most trusted part of the system. The realization of these primitives on the underlying processing platform plays a crucial role for any real-world deployment. In this work, we discuss new trends in public-key cryptography that could potentially establish as alternatives to the currently used RSA and ECC cryptosystems. Analyzing these trends from a developer's perspective, we identify the requirements for optimal processing architectures. Moreover, we investigate if these requirements are already satisfied with latest processing platforms, i.e., which of the streaming, hybrid, multi-core or application-specific instructions processors provide a most promising target architecture.
Tim Güneysu, Stefan Heyse, Christof Paar
ACM Great Lakes Symposium on VLSI1
2010 Breaking Elliptic Curve Cryptosystems Using Reconfigurable Hardware
abstract
This paper reports a new speed record for FPGAs in cracking Elliptic Curve Cryptosystems. We conduct a detailed analysis of different F2(m)multiplication approaches in this application. A novel architecture using optimized normal basis multipliers is proposed to solve the Certicom challenge ECC2K-130. We compare the FPGA performance against CPUs, GPUs, and the Sony PlayStation 3. Our implementations show low-cost FPGAs outperform even multicore desktop processors and graphics cards by a factor of 2.
Junfeng Fan, Daniel V. Bailey, Lejla Batina, Tim Güneysu, Christof Paar, Ingrid Verbauwhede
FPL4
2010 High-Performance Integer Factoring with Reconfigurable Devices
abstract
We present a novel FPGA-based implementation of the Elliptic Curve Method (ECM) for the factorization of medium-sized composite integers. More precisely, we demonstrate an ECM implementation capable to determine prime factors of up to 2,424 151-bit integers per second using a single Xilinx Virtex-4 SX35 FPGA. Using this implementation on a cluster like the COPACOBANA is beneficial for attacking cryptographic primitives like the well-known RSA cryptosystem with advanced methods such as the Number Field Sieve (NFS). To provide this vast number of integer factorizations per FPGA, we make use of the available DSP blocks on each Virtex-4 device to accelerate low-level arithmetic computations. This methodology allows the development of a time-area efficient design that runs 24 ECM cores in parallel, implementing both phase 1 and phase 2 of the ECM. Moreover, our design is fully scalable and supports composite integers in the range from 66 to 236 bits without any significant modifications to the hardware. Compared to the implementation by Gaj et al., who reported an ECM design for the same Virtex-4 platform, our improved architecture provides an advanced cost-performance ratio which is better by a factor of 37.
Ralf Zimmermann 0001, Tim Güneysu, Christof Paar
FPL2
2010 True random number generation in block memories of reconfigurable devices
abstract
We present concepts and implementations to transform write collisions in memory blocks into an entropy source for random number generation. Write collisions in dual-ported block memories occur when both memory ports write simultaneously different data at the same memory location. After a thorough analysis of this effect, we present a robust methodology to generate digitized noise and randomness from such write collisions and also provide details how to implement post-processing methods for efficient bias and correlation removal. Finally, we present three concepts and implementations for random number generators stages that can deliver random data at an output rate of more than 100 MBit/s.
Tim Güneysu
FPT1
2010 DSPs, BRAMs, and a Pinch of Logic: Extended Recipes for AES on FPGAs
abstract
We present three lookup-table-based AES implementations that efficiently use the BlockRAM and DSP units embedded within Xilinx Virtex-5 FPGAs. An iterative module outputs a 32-bit AES round column every clock cycle, with a throughput of 1.67 Gbit/s when processing two 128-bit inputs. This construct is then replicated four times to provide a complete AES round per cycle with 6.7 Gbit/s throughput when processing eight input streams. This, in turn, is replicated ten times for a fully unrolled design providing over 52 Gbit/s of throughput. We also present implementations of a BRAM-based AES key-expansion, CMAC, and CTR modes of operation. Results for designs where DSPs are replaced by regular logic are also presented. The combination and arrangement of the specialized embedded functions available in the FPGA allows us to implement our designs using very few traditional user logic elements such as flip-flops and lookup tables, yet still achieve these high throughputs. HDL source code, simulation testbenches, and software tool commands to reproduce reported results for the three AES variants and CMAC mode are made publicly available. Our contribution concludes with a discussion on comparing cipher implementations in the literature, and why these comparisons can be meaningless without a common reporting methodology, or within the context of a constrained target application.
Saar Drimer, Tim Güneysu, Christof Paar
ACM Trans. Reconfigurable Technol. Syst.2
2009 MicroEliece: McEliece for Embedded Devices
Thomas Eisenbarth 0001, Tim Güneysu, Stefan Heyse, Christof Paar
CHES2
2009 Trojan Side-Channels: Lightweight Hardware Trojans through Side-Channel Engineering
Lang Lin, Markus Kasper, Tim Güneysu, Christof Paar, Wayne P. Burleson
CHES3
2009 Transforming write collisions in block RAMs into security applications
abstract
Due to their versatile and generic structure, field programmable gate arrays (FPGA) allow dynamic reconfiguration of their logical resources just by loading configuration files. However, this flexibility also opens up the threat of theft of intellectual property (IP) since these configuration files can be easily extracted and cloned. In this context, the ability to bind a configuration to a specific device is an important step to prevent product counterfeiting. In this paper, we present a novel strategy to identify and authenticate FPGAs in applications using intrinsic, device-specific information (also known as physically unclonable functions). Our solution is based on the output of intentionally induced write collisions in synchronous dual-port block RAM (BRAM). We show that the output of such write collisions can be used to create unique device signatures. In addition to applications for chip identification and authentication, we also propose a solution to efficiently create secret keys on-chip. As a last contribution, we outline how to transform our idea into a circuit for true random number generation (TRNG).
Tim Güneysu, Christof Paar
FPT1
2008 Ultra High Performance ECC over NIST Primes on Commercial FPGAs
Tim Güneysu, Christof Paar
CHES1
2008 Exploiting the Power of GPUs for Asymmetric Cryptography
Robert Szerwinski, Tim Güneysu
CHES2
2008 DSPs, BRAMs and a Pinch of Logic: New Recipes for AES on FPGAs
abstract
We present an AES cipher implementation that is based on the BlockRAM and DSP units embedded within Xilinx's Virtex-5 FPGAs. An iterative "basic" module outputs a 32 bit column of an AES round each clock cycle, with a throughput of 1.76 Gbit/s when processing two 128 bit inputs. This construct is replicated four times for a 128 bit datapath for a full AES round with 6.21 Gbit/s throughput when processing eight inputs. Finally, the "round" module is replicated ten times for a fully unrolled design that yields over 55 Gbit/s of throughput. The combination and arrangement of the specialized embedded functions available in the FPGA allows us to implement our designs using very few traditional user logic elements such as flip-flops and lookup tables, yet still achieve these high throughputs. The complete source code for these designs is made publicly available for use in further research and for replicating our results. Our contribution ends with a discussion of comparing cipher implementations in the literature, and why these comparisons can be meaningless without a common reporting style, platform, or within the context of a specific constrained application.
Saar Drimer, Tim Güneysu, Christof Paar
FCCM2
2008 Enhancing COPACOBANA for advanced applications in cryptography and cryptanalysis
abstract
Cryptanalysis of symmetric and asymmetric ciphers is a challenging task due to the enormous amount of involved computations. To tackle this computational complexity, usually the employment of special-purpose hardware is considered as best approach. We have built a massively parallel cluster system (COPACOBANA) based on low-cost FPGAs as a cost-efficient platform primarily targeting cryptanalytical operations with these high computational efforts but low communication and memory requirements. However, some parallel applications in the field of cryptography are too complex for low-cost FPGAs and also require the availability of at least moderate communication and memory facilities. Particularly, this holds true for arithmetic intensive application as well as ones with a highly complex data flow. In this contribution, we describe a novel architecture for a more versatile and reliable COPACOBANA capable to host advanced cryptographic applications like high-performance digital signature generation according to the elliptic curve digital signature algorithm (ECDSA) and integer factorization based on the elliptic curve method (ECM). In addition to that, the new cluster design allows even to run more supercomputing applications beyond the field of cryptography.
Tim Güneysu, Christof Paar, Gerd Pfeiffer, Manfred Schimmler
FPL1
2008 Cryptanalysis with COPACOBANA
abstract
Cryptanalysis of ciphers usually involves massive computations. The security parameters of cryptographic algorithms are commonly chosen so that attacks are infeasible with available computing resources. This contribution presents a variety of cryptanalytical applications utilizing the COPACOBANA (Cost-Optimized Parallel Code Breaker) machine which is a high-performance, low-cost cluster consisting of 120 Field Programmable Gate Arrays (FPGA). COPACOBANA appears to be the only such reconfigurable parallel FPGA machine optimized for code breaking tasks reported in the open literature. Depending on the actual algorithm, the parallel hardware architecture can outperform conventional computers by several orders of magnitude. In this work, we will focus on novel implementations of cryptanalytical algorithms, utilizing the impressive computational power of COPACOBANA. We describe various exhaustive key search attacks on symmetric ciphers and demonstrate an attack on a security mechanism employed in the electronic passport. Furthermore, we describe time-memory tradeoff techniques which can, e.g., be used for attacking the popular A5/1 algorithm used in GSM voice encryption. In addition, we introduce efficient implementations of more complex cryptanalysis on asymmetric cryptosystems, e.g., Elliptic Curve Cryptosystems (ECC) and number co-factorization for RSA.
Tim Güneysu, Timo Kasper, Martin Novotný, Christof Paar, Andy Rupp
IEEE Trans. Computers1
2008 Special-Purpose Hardware for Solving the Elliptic Curve Discrete Logarithm Problem
abstract
The resistance against powerful index-calculus attacks makes Elliptic Curve Cryptosystems (ECC) an interesting alternative to conventional asymmetric cryptosystems, like RSA. Operands in ECC require significantly less bits at the same level of security, resulting in a higher computational efficiency compared to RSA. With growing computational capabilities and continuous technological improvements over the years, however, the question of the security of ECC against attacks based on special-purpose hardware arises. In this context, recently emerged low-cost FPGAs demand for attention in the domain of hardware-based cryptanalysis: the extraordinary efficiency of modern programmable hardware devices allow for a low-budget implementation of hardware-based ECC attacks---without the requirement of the expensive development of ASICs. With focus on the aspect of cost-efficiency, this contribution presents and analyzes an FPGA-based architecture of an attack against ECC over prime fields. A multi-processing hardware architecture for Pollard's Rho method is described. We provide results on actually used key lengths of ECC (128 bits and above) and estimate the expected runtime for a successful attack. As a first result, currently used elliptic curve cryptosystems with a security of 160 bit and above turn out to be infeasible to break with available computational and financial resources. However, some of the security standards proposed by the Standards for Efficient Cryptography Group (SECG) become subject to attacks based on low-cost FPGAs.
Tim Güneysu, Christof Paar, Jan Pelzl
ACM Trans. Reconfigurable Technol. Syst.1
2007 Establishing Chain of Trust in Reconfigurable Hardware
abstract
Facing ubiquitous threats like computer viruses, trojans and theft of intellectual property, Trusted computing (TC) is an emerging technology towards building trustworthy computing platforms. A recent initiative by the trusted computing group (TCG) specifies the use of trusted platform modules (TPM), currently implemented as dedicated, cost-effective crypto-chips mounted on the main board of computer systems. In this paper we propose implementations for TC functionalities based on more flexible and versatile approaches for reconfigurable and embedded architectures. Our approach allows for (i) a scalable design and update of TPM functionalities in embedded systems, (ii) the integration of the TPM hardware in the chain of trust to bind applications to the underlying TPM and the reconfigurable hardware, and (iii) the design of vendor independent TPMs.
Thomas Eisenbarth 0001, Tim Güneysu, Christof Paar, Ahmad-Reza Sadeghi, Marko Wolf, Russell Tessier
FCCM2
2007 New Protection Mechanisms for Intellectual Property in Reconfigurable Logic
abstract
The distinct advantage of SRAM-based Field Programmable Gate Arrays (FPGA) is their flexibility for configuration changes. But this opens up the threat of intellectual property (IP) theft since the system configuration is stored in easy-to-access Flash memory. High-end FPGAs have already been extended with symmetric-key decryption engines used to load an encrypted version of the configuration that cannot simply be copied and used without knowledge of the secret key. However, with respect to business and licensing processes, this protection system lacks a convenient scheme for key transport and installation. We propose a new protection scheme for the IP of circuits in configuration bit files that provides a significant improvement to the current unsatisfying situation. It uses both public-key and symmetric cryptography, but does not burden FPGAs with the usual overhead of public-key cryptography: While it needs hard-wired symmetric cryptography, the public-key functionality is moved into a temporary configuration bit stream for a one-time setup procedure. This approach requires only very few modifications to current FPGA technology. Using five basic stages, the new protection scheme allows new accounting models for volume licensing of IP, with automated key installation on FPGAs taking place at the customer's site.
Tim Güneysu, Bodo Möller, Christof Paar
FCCM1
2007 Attacking elliptic curve cryptosystems with special-purpose hardware
abstract
Since their invention in the mid 1980s, Elliptic Curve Cryptosystems (ECC) have become an alternative to common Public-Key (PK) cryptosystems such as, e.g., RSA. The utilization of Elliptic Curves (EC) in cryptography is very promising because of their resistance against powerful index-calculus attacks. Providing a similar level of security as RSA, ECC allows for efficient implementation due to a significantly smaller bit size of the operands. It is widely accepted that the only feasible way to attack actual cryptosystems, if at all, is the application of dedicated hardware. In times of continuous technological improvements and increasing computing power, the question of the security of ECC against attacks based on special-purpose hardware and, in particular based on recently emerged low-cost FPGAs, arises.This work presents the first architecture with a corresponding FPGA implementation of an attack against ECC over prime fields. We describe an FPGA-based multi-processing hardware architecture for the Pollard-Rho method which is, to our knowledge, currently the most efficient attack against ECC. The implementation is running on a contemporary low-cost FPGA which allows for a much better cost-performance ratio than conventional CPUs. With the implementation at hand, a fairly accurate estimate about the cost of an FPGA-based attack can be given. We will extrapolate the results on actual ECC key lengths (128 bits and above) and estimate the expected runtimes for a successful attack. Since FPGA-based attacks are out of reach for key lengths exceeding 128 bits, we provide estimates for an ASIC design.Based on our results, currently used elliptic curve cryptosystems (160 bit and above) are infeasible to break with available computational and financial resources. However, some of the security standards proposed by the SECG in [2, 3] become subject to attacks based on low-cost FPGAs.
Tim Güneysu, Christof Paar, Jan Pelzl
FPGA1
2007 Dynamic Intellectual Property Protection for Reconfigurable Devices
abstract
The distinct advantage of SRAM-based field programmable gate arrays (FPGA) is their flexibility for configuration changes. However, this opens up the threat of theft of Intellectual Property (IP) since the system configuration is stored in easy-to-access Flash memory. To prevent this, high-end FPGAs have already been extended with symmetric-key decryption engines used to load an encrypted version of the configuration that cannot simply be copied and used without knowledge of the secret key. However, such protection systems based on straightforward use of symmetric cryptography are not well-suited with respect to business and licensing processes, since they are lacking a convenient scheme for key transport and installation. We propose a new protection scheme for the IP of circuits in configuration bit files that provides a significant improvement to the current unsatisfying situation. It uses both public-key and symmetric cryptography, but does not burden FPGAs with the usual overhead of public-key cryptography: While it needs hard-wired symmetric cryptography, the public-key functionality is moved into a temporary configuration bit stream for a one-time setup procedure. This approach requires only very few modifications to current FPGA technology. Using five basic stages, the new protection scheme allows new accounting models for volume licensing of IP, with automated key installation on FPGAs taking place at the customer's site.
Tim Güneysu, Bodo Möller, Christof Paar
FPT1