Johann Großschädl

dblp:g/JGrossschadl · DBLP profile ↗
← Back
64ranked-venue papers
19as first author
6since 2021 · last 2025
0009-0006-3210-3102ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 48 · 12 first-author · 3 since 2021Systems, architecture and hardware · 11 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Theory of computation · 2Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 High-Throughput EdDSA Verification on Intel Processors with Advanced Vector Extensions
Hao Cheng 0009, Johann Großschädl, Peter Y. A. Ryan
SAC3
2024 RISC-V Instruction Set Extensions for Multi-Precision Integer Arithmetic: A Case Study on Post-Quantum Key Exchange Using CSIDH-512
abstract
Multi-Precision Integer (MPI) arithmetic is a performance-critical component of many public-key cryptosystems, including besides classical ones (e.g., RSA, ECC) also isogeny-based post-quantum schemes. In this paper, we analyze and compare two widely-used MPI representations, namely full-radix and reduced-radix, for the efficient implementation of modular arithmetic operations on the 64-bit RISC-V (RV64GC) architecture. We also evaluate how the execution times of both can be further improved with Instruction Set Extensions (ISEs). The ISEs we propose are able to accelerate a CSIDH-512 class group action by a factor of 1.71 compared to a standard software implementation on a 64-bit Rocket core. This speed-up comes at the cost of a hardware overhead of about 10%.
Hao Cheng 0009, Georgios Fotiadis, Johann Großschädl, Dan Page, Thinh Hung Pham, Peter Y. A. Ryan
DAC3
2022 Rivain-Prouff on Steroids: Faster and Stronger Masking of the AES
Luan Cardoso dos Santos, François Gérard, Johann Großschädl, Lorenzo Spignoli
CARDIS3
2022 Efficient Software Implementation of the SIKE Protocol Using a New Data Representation
abstract
Thanks to relatively small public and secret keys, the Supersingular Isogeny Key Encapsulation (SIKE) protocol made it into the third evaluation round of the post-quantum standardization project of the National Institute of Standards and Technology (NIST). Even though a large body of research has been devoted to the efficient implementation of SIKE, its latency is still undesirably long for many real-world applications. Most existing implementations of the SIKE protocol use the Montgomery representation for the underlying field arithmetic since the corresponding reduction algorithm is considered the fastest method for performing multiple-precision modular reduction. In this paper, we propose a new data representation for supersingular isogeny-based Elliptic-Curve Cryptography (ECC), of which SIKE is a sub-class. This new representation enables significantly faster implementations of modular reduction than the Montgomery reduction, and also other finite-field arithmetic operations used in ECC can benefit from our data representation. We implemented all arithmetic operations in C using the proposed representation such that they have constant execution time and integrated them to the latest version of the SIKE software library. Using four different parameters sets, we benchmarked our design and the optimized generic implementation on a 2.6 GHz Intel Xeon E5-2690 processor. Our results show that, for the prime of SIKEp751, the proposed reduction algorithm is approximately 2.61 times faster than the currently best implementation of Montgomery reduction, and our representation also enables significantly better timings for other finite-field operations. Due to these improvements, we were able to achieve a speed-up by a factor of about 1.65, 2.03, 1.61, and 1.48 for SIKEp751, SIKEp610, SIKEp503, and SIKEp434, respectively, compared to state-of-the-art generic implementations.
Jing Tian 0004, Piaoyang Wang, Zhe Liu 0001, Jun Lin 0001, Zhongfeng Wang 0001, Johann Großschädl
IEEE Trans. Computers6
2021 AVRNTRU: Lightweight NTRU-based Post-Quantum Cryptography for 8-bit AVR Microcontrollers
abstract
Introduced in 1996, NTRUEncrypt is not only one of the earliest but also one of the most scrutinized lattice-based cryptosystems and expected to remain secure in the upcoming era of quantum computing. Furthermore, NTRUEncrypt offers some efficiency benefits over “pre-quantum” cryptosystems like RSA or ECC since the low-level arithmetic operations are less computation-intensive and, thus, more suitable for constrained devices. In this paper we present Avrntru, a highly-optimized implementation of NTRUEncrypt for 8-bit AVR microcontrollers that we developed from scratch to reach high performance and resistance to timing attacks. Avrntru complies with the EESS #1 v3.1 specification and supports product-form parameter sets such as ees443ep1, ees587ep1, and ees743ep1. An entire encryption (including mask generation and blinding-polynomial generation) using the ees443ep1 parameters requires 847973 clock cycles on an ATmega1281 microcontroller; the decryption is more costly and has an execution time of 1051871 cycles. We achieved these results with the help of a novel hybrid technique for multiplication in a truncated polynomial ring, whereby one of the operands is a sparse ternary polynomial in product form and the other an arbitrary element of the ring. A constant-time multiplication in the ring given by the ees443ep1 parameters takes only 192577 cycles, which sets a new speed record for the arithmetic part of a lattice-based cryptosystem on AVR.
Hao Cheng 0009, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
DATE2
2021 Lightweight EdDSA Signature Verification for the Ultra-Low-Power Internet of Things
Johann Großschädl, Christian Franck, Zhe Liu 0001
ISPEC1
2020 Parallel Implementation of SM2 Elliptic Curve Cryptography on Intel Processors with AVX2
Junhao Huang 0001, Zhe Liu 0001, Johann Großschädl
ACISP4
2020 Lightweight Post-quantum Key Encapsulation for 8-bit AVR Microcontrollers
Hao Cheng 0009, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
CARDIS2
2020 Alzette: A 64-Bit ARX-box - (Feat. CRAX and TRAX)
Christof Beierle, Alex Biryukov, Luan Cardoso dos Santos, Johann Großschädl, Léo Perrin, Aleksei Udovenko, Vesselin Velichkov, Qingju Wang 0001
CRYPTO (3)4
2020 High-Throughput Elliptic Curve Cryptography Using AVX2 Vector Instructions
Hao Cheng 0009, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
SAC2
2019 FELICS-AEAD: Benchmarking of Lightweight Authenticated Encryption Algorithms
Luan Cardoso dos Santos, Johann Großschädl, Alex Biryukov
CARDIS2
2019 A Lightweight Implementation of NTRU Prime for the Post-quantum Internet of Things
Hao Cheng 0009, Daniel Dinu, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
WISTP3
2018 A Family of Lightweight Twisted Edwards Curves for the Internet of Things
Sankalp Ghatpande, Johann Großschädl, Zhe Liu 0001
WISTP2
2017 Efficient Masking of ARX-Based Block Ciphers Using Carry-Save Addition on Boolean Shares
Daniel Dinu, Johann Großschädl, Yann Le Corre
ISC2
2017 Elliptic Curve Cryptography with Efficiently Computable Endomorphisms and Its Hardware Implementations for the Internet of Things
abstract
Verification of an ECDSA signature requires a double scalar multiplication on an elliptic curve. In this work, we study the computation of this operation on a twisted Edwards curve with an efficiently computable endomorphism, which allows reducing the number of point doublings by approximately 50 percent compared to a conventional implementation. In particular, we focus on a curve defined over the 207-bit prime field Fpwith p = 2207- 5,131. We develop several optimizations to the operation and we describe two hardware architectures for computing the operation. The first architecture is a small processor implemented in 0.13 μm CMOS ASIC and is useful in resource-constrained devices for the Internet of Things (IoT) applications. The second architecture is designed for fast signature verifications by using FPGA acceleration and can be used in the server-side of these applications. Our designs offer various trade-offs and optimizations between performance and resource requirements and they are valuable for IoT applications.
Zhe Liu 0001, Johann Großschädl, Kimmo Järvinen 0001, Husen Wang, Ingrid Verbauwhede
IEEE Trans. Computers2
2017 High-Performance Ideal Lattice-Based Cryptography on 8-Bit AVR Microcontrollers
abstract
Over recent years lattice-based cryptography has received much attention due to versatile average-case problems like Ring-LWE or Ring-SIS that appear to be intractable by quantum computers. In this work, we evaluate and compare implementations of Ring-LWE encryption and the bimodal lattice signature scheme (BLISS) on an 8-bit Atmel ATxmega128 microcontroller. Our implementation of Ring-LWE encryption provides comprehensive protection against timing side-channels and takes 24.9ms for encryption and 6.7ms for decryption. To compute a BLISS signature, our software takes 317ms and 86ms for verification. These results underline the feasibility of lattice-based cryptography on constrained devices.
Zhe Liu 0001, Thomas Pöppelmann, Tobias Oder, Hwajeong Seo, Sujoy Sinha Roy, Tim Güneysu, Johann Großschädl, Howon Kim 0001, Ingrid Verbauwhede
ACM Trans. Embed. Comput. Syst.7
2016 Energy-Efficient Elliptic Curve Cryptography for MSP430-Based Wireless Sensor Nodes
Zhe Liu 0001, Johann Großschädl, Lin Li 0041, Qiuliang Xu
ACISP (1)2
2016 Correlation Power Analysis of Lightweight Block Ciphers: From Theory to Practice
Alex Biryukov, Daniel Dinu, Johann Großschädl
ACNS3
2016 Design Strategies for ARX with Provable Bounds: Sparx and LAX
abstract
We present, for the first time, a general strategy for designing ARX symmetric-key primitives with provable resistance against single-trail differential and linear cryptanalysis. The latter has been a long standing open problem in the area of ARX design. The wide-trail design strategy (WTS), that is at the basis of many S-box based ciphers, including the AES, is not suitable for ARX designs due to the lack of S-boxes in the latter. In this paper we address the mentioned limitation by proposing the long trail design strategy (LTS) – a dual of the WTS that is applicable (but not limited) to ARX constructions. In contrast to the WTS, that prescribes the use of small and efficient S-boxes at the expense of heavy linear layers with strong mixing properties, the LTS advocates the use of large (ARX-based) S-Boxes together with sparse linear layers. With the help of the so-called long-trail argument , a designer can bound the maximum differential and linear probabilities for any number of rounds of a cipher built according to the LTS. To illustrate the effectiveness of the new strategy, we propose Sparx – a family of ARX-based block ciphers designed according to the LTS. Sparx has 32-bit ARX-based S-boxes and has provable bounds against differential and linear cryptanalysis. In addition, Sparx is very efficient on a number of embedded platforms. Its optimized software implementation ranks in the top 6 of the most software-efficient ciphers along with Simon , Speck , Chaskey, LEA and RECTANGLE. As a second contribution we propose another strategy for designing ARX ciphers with provable properties, that is completely independent of the LTS. It is motivated by a challenge proposed earlier by Wallén and uses the differential properties of modular addition to minimize the maximum differential probability across multiple rounds of a cipher. A new primitive, called LAX , is designed following those principles. LAX partly solves the Wallén challenge.
Daniel Dinu, Léo Perrin, Aleksei Udovenko, Vesselin Velichkov, Johann Großschädl, Alex Biryukov
ASIACRYPT (1)5
2016 Efficient arithmetic on ARM-NEON and its application for high-speed RSA implementation
abstract
Abstract Advanced modern processors support single instruction, multiple data instructions (e.g., Intel‐AVX and ARM‐NEON) and a massive body of research on vector‐parallel implementations of modular arithmetic, which are crucial components for modern public‐key cryptography ranging from Rivest, Shamir, and Adleman (RSA), ElGamal, Digital Signature Algorithm (DSA), and elliptic curve cryptography, have been conducted. In this paper, we introduce a novel double operand scanning method to speed up multi‐precision squaring with non‐redundant representations on single instruction, multiple data architecture where the part of the operands are doubled to compute the squaring operation without read‐after‐write dependencies between source and destination variables. Afterwards, Karatsuba algorithm is applied to both multiplication and squaring operations. For modular multiplication, separated Montgomery algorithm is chosen. Finally, the Rivest, Shamir, and Adleman (RSA) implementations outperform the best‐known results on the ARM‐NEON platforms. Copyright © 2017 John Wiley & Sons, Ltd.
Hwajeong Seo, Zhe Liu 0001, Johann Großschädl, Howon Kim 0001
Secur. Commun. Networks3
2016 Efficient Implementation of NIST-Compliant Elliptic Curve Cryptography for 8-bit AVR-Based Sensor Nodes
abstract
In this paper, we introduce a highly optimized software implementation of standards-compliant elliptic curve cryptography (ECC) for wireless sensor nodes equipped with an 8-bit AVR microcontroller. We exploit the state-of-the-art optimizations and propose novel techniques to further push the performance envelope of a scalar multiplication on the NIST P-192 curve. To illustrate the performance of our ECC software, we develope the prototype implementations of different cryptographic schemes for securing communication in a wireless sensor network, including elliptic curve Diffie–Hellman (ECDH) key exchange, the elliptic curve digital signature algorithm (ECDSA), and the elliptic curve Menezes–Qu–Vanstone (ECMQV) protocol. We obtain record-setting execution times for fixed-base, point variable-base, and double-base scalar multiplication. Compared with the related work, our ECDH key exchange achieves a performance gain of roughly 27% over the best previously published result using the NIST P-192 curve on the same platform, while our ECDSA performs twice as fast as the ECDSA implementation of the well-known TinyECC library. We also evaluate the impact of Karatsuba’s multiplication technique on the overall execution time of a scalar multiplication. In addition to offering high performance, our implementation of scalar multiplication has a highly regular execution profile, which helps to protect against certain side-channel attacks. Our results show that NIST-compliant ECC can be implemented efficiently enough to be suitable for resource-constrained sensor nodes.
Zhe Liu 0001, Hwajeong Seo, Johann Großschädl, Howon Kim 0001
IEEE Trans. Inf. Forensics Secur.3
2015 Efficient Implementation of ECDH Key Exchange for MSP430-Based Wireless Sensor Networks
abstract
Public-Key Cryptography (PKC) is an indispensable building block of modern security protocols, and, thus, essential for secure communication over insecure networks. Despite a significant body of research devoted to making PKC more "lightweight," it is still commonly perceived that software implementations of PKC are computationally too expensive for practical use in ultra-low power devices such as wireless sensor nodes. In the present paper we aim to challenge this perception and present a highly-optimized implementation of Elliptic Curve Cryptography (ECC) for the TI MSP430 series of 16-bit microcontrollers. Our software is inspired by MoTE-ECC and supports scalar multiplication on two families of elliptic curves, namely Montgomery and twisted Edwards curves. However, in contrast to MoTE-ECC, we use pseudo-Mersenne prime fields as underlying algebraic structure to facilitate inter-operability with existing ECC implementations. We introduce a novel "zig-zag" technique for multiple-precision squaring on the MSP430 and assess its execution time. Similar to MoTE-ECC, we employ the Montgomery model for variable-base scalar multiplications and the twisted Edwards model if the base point is fixed (e.g. to generate an ephemeral key pair). Our experiments show that the two scalar multiplications needed to perform an ephemeral ECDH key exchange can be accomplished in 4.88 million clock cycles altogether (using a 159-bit prime field), which sets a new speed record for ephemeral ECDH on a 16-bit processor. We also describe the curve generation process and analyze the execution time of various field and point arithmetic operations on curves over a 159-bit and a 191-bit pseudo-Mersenne prime field.
Zhe Liu 0001, Hwajeong Seo, Xinyi Huang 0001, Johann Großschädl
AsiaCCS5
2015 Efficient Ring-LWE Encryption on 8-Bit AVR Processors
Zhe Liu 0001, Hwajeong Seo, Sujoy Sinha Roy, Johann Großschädl, Howon Kim 0001, Ingrid Verbauwhede
CHES4
2015 Higher-Order Masking in Practice: A Vector Implementation of Masked AES for ARM NEON
Junwei Wang 0003, Praveen Kumar Vadnala, Johann Großschädl, Qiuliang Xu
CT-RSA3
2015 Conversion from Arithmetic to Boolean Masking with Logarithmic Complexity
Jean-Sébastien Coron, Johann Großschädl, Mehdi Tibouchi, Praveen Kumar Vadnala
FSE2
2014 MoTE-ECC: Energy-Scalable Elliptic Curve Cryptography for Wireless Sensor Networks
Zhe Liu 0001, Erich Wenger, Johann Großschädl
ACNS3
2014 Secure Conversion between Boolean and Arithmetic Masking of Any Order
Jean-Sébastien Coron, Johann Großschädl, Praveen Kumar Vadnala
CHES2
2014 Reverse Product-Scanning Multiplication and Squaring on 8-Bit AVR Processors
Zhe Liu 0001, Hwajeong Seo, Johann Großschädl, Howon Kim 0001
ICICS3
2014 High-Speed Elliptic Curve Cryptography on the NVIDIA GT200 Graphics Processing Unit
Shujie Cui, Johann Großschädl, Zhe Liu 0001, Qiuliang Xu
ISPEC2
2014 Design and implementation of a versatile cryptographic unit for RISC processors
abstract
In this paper, we design, implement, and realize a cryptographic unit (CU) that can easily be integrated to any reduced instruction set computing (RISC)-type processor for the safe and efficient execution of cryptographic algorithms. Design of the CU takes a novel approach in the execution of cryptographic algorithms when compared with cryptographic accelerators and architectural enhancements. Although it is integrated to a pipeline of an embedded RISC processor, it is partially an autonomous unit with its own resources, which is analogous to the floating point unit in this sense. It provides new instructions to accelerate cryptographic algorithms, and its associated cost in terms of area is acceptable and justified by the improvement in the performance and efficiency. The CU can also be instrumental in protecting the cryptographic computation against active and passive attacks and other malicious processes running simultaneously. We demonstrate that the execution of Advanced Encryption Standart (AES) encryption can be performed inside the CU, which prevents secret and/or sensitive information from leaving the CU during the cryptographic computation.
Kazim Yumbul, Erkay Savas, Övünç Kocabas, Johann Großschädl
Secur. Commun. Networks4
2013 Low-Weight Primes for Lightweight Elliptic Curve Cryptography on 8-bit AVR Processors
Zhe Liu 0001, Johann Großschädl, Duncan S. Wong
Inscrypt2
2013 Efficient Implementation of NIST-Compliant Elliptic Curve Cryptography for Sensor Nodes
Zhe Liu 0001, Hwajeong Seo, Johann Großschädl, Howon Kim 0001
ICICS3
2012 Efficient Java Implementation of Elliptic Curve Cryptography for J2ME-Enabled Mobile Devices
Johann Großschädl, Dan Page, Stefan Tillich
WISTP1
2012 Cryptanalysis of the Full AES Using GPU-Like Special-Purpose Hardware
abstract
The block cipher Rijndael has undergone more than ten years of extensive cryptanalysis since its submission as a candidate for the Advanced Encryption Standard (AES) in April 1998. To date, most of the publicly-known cryptanalytic results are based o
Alex Biryukov, Johann Großschädl
Fundam. Informaticae2
2011 An Exploration of Mechanisms for Dynamic Cryptographic Instruction Set Extension
Philipp Grabher, Johann Großschädl, Simon Hoerder, Kimmo Järvinen 0001, Dan Page, Stefan Tillich, Marcin Wójcik
CHES2
2011 A Unified Multiply/Accumulate Unit for Pairing-Based Cryptography over Prime, Binary and Ternary Fields
abstract
Bilinear maps, or pairings, on elliptic curves are an active area of research in modern cryptology with applications ranging from cryptanalysis (e.g. MOV attack) over identity-based encryption to short signature schemes. Many parameterisations and implementation options for pairing-based cryptography have been investigated in the recent past. Elliptic curves over prime fields are often preferred for software implementation, whereas extension fields of characteristic two and three offer advantages for implementation in hardware. In the ideal case, a hardware accelerator for pairing-based cryptography can support all three types of field to ensure inter-operability with a broad spectrum of applications. This need has motivated the design of so-called unified multipliers, which are basically multipliers that integrate different types of operands (e.g. integers and polynomials) into a single data path. In the present paper, we introduce a unified multiply/accumulate unit for signed/unsigned integers as well as binary and ternary polynomials. The multiplier generates partial products using a Redundant Signed-Digit (RSD) representation that allows for efficient combination of all three operand types into one data path. In addition, our design takes advantage of a high-radix encoding scheme for integers and binary polynomials to reduce the overall number of partial products and utilise the data path in an optimal way. We compare our multiplier with a previous radix-2 implementation of Ozturk et al and analyse the differences in terms of silicon area and critical path delay. The unified multiply/accumulate unit described in this paper can be used in embedded systems like smart cards, either as arithmetic core of a cryptographic co-processor, or as functional unit of an application-specific processor.
Tobias Vejda, Johann Großschädl, Dan Page
DSD2
2010 Performance and Security Aspects of Client-Side SSL/TLS Processing on Mobile Devices
Johann Großschädl, Ilya Kizhvatov
CANS1
2009 Full-Custom VLSI Design of a Unified Multiplier for Elliptic Curve Cryptography on RFID Tags
Johann Großschädl
Inscrypt1
2009 Hardware/Software Co-design of Public-Key Cryptography for SSL Protocol Execution in Embedded Systems
Manuel Koschuch, Johann Großschädl, Dan Page, Philipp Grabher, Matthias Hudler, Michael Krüger
ICICS2
2009 Energy-Efficient Implementation of ECDH Key Exchange for Wireless Sensor Networks
Christian Lederer, Roland Mader, Manuel Koschuch, Johann Großschädl, Alexander Szekely, Stefan Tillich
WISTP4
2008 Workload Characterization of a Lightweight SSL Implementation Resistant to Side-Channel Attacks
Manuel Koschuch, Johann Großschädl, Udo Payer, Matthias Hudler, Michael Krüger
CANS2
2008 Light-Weight Instruction Set Extensions for Bit-Sliced Cryptography
Philipp Grabher, Johann Großschädl, Dan Page
CHES2
2007 The energy cost of cryptographic key establishment in wireless sensor networks
abstract
Wireless sensor nodes generally face serious limitations in terms of computational power, energy supply, and network bandwidth. Therefore, the implementation of effective and secure techniques for setting up a shared secret key between sensor nodes is a challenging task. In this paper we analyze and compare the energy cost of two different protocols for authenticated key establishment. The first protocol employs a lightweight variant of the Kerberos key transport mechanism with 128-bit AES encryption. The second protocol is based on ECMQV, an authenticated version of the elliptic curve Diffie-Hellman key exchange, and uses a 256-bit prime field GF(p) as underlying algebraic structure. We evaluate the energy cost of both protocols on a Rockwell WINS node equipped with a 133 MHz Strong ARM processor and a 100 kbit/s radio module. The evaluation considers both the processor's energy consumption for calculating cryptographic primitives and the energy cost of radio communication for different transmit power levels. Our simulation results show that the ECMQV key exchange consumes up to twice as much energy as Kerberos-like key transport.
Johann Großschädl, Alexander Szekely, Stefan Tillich
AsiaCCS1
2007 Power Analysis Resistant AES Implementation with Instruction Set Extensions
Stefan Tillich, Johann Großschädl
CHES2
2007 Energy evaluation of software implementations of block ciphers under memory constraints
Johann Großschädl, Stefan Tillich, Christian Rechberger, Michael Hofmann 0007, Marcel Medwed
DATE1
2007 Performance Evaluation of Instruction Set Extensions for Long Integer Modular Arithmetic on a SPARC V8 Processor
abstract
Many important algorithms for public-key cryptography rely on computation-intensive arithmetic operations like modular exponentiation on very long integers, typically in the range of 512 and 2048 bits. Modular exponentiation is generally realized through a sequence of modular multiplications and spends the majority of execution time in simple inner loops. Speeding up these performance-critical inner loop operations with custom instructions has, therefore, a significant impact on the total execution time of public-key cryptosystems. In this paper we analyze the performance of instruction set extensions for long integer arithmetic on a SPARC V8 processor. We discuss various implementation options and optimization opportunities for both modular multiplication and exponentiation. In particular, we introduce a partial loop unrolling (PLU) technique for modular multiplication which allows to achieve large performance gains at the cost of a moderate increase in code size, while maintaining the full flexibility of a "rolled-loop" implementation. In addition, we study window methods for modular exponentiation and analyze their impact on performance and memory requirements. Our experimental results, obtained with an FPGA prototype of the LEON-2 SPARC V8 core, show that a full 1024-bit modular exponentiation can be performed in about 12.5 ldr 106clock cycles, which is a reasonable value for embedded devices like smart cards or sensor nodes.
Johann Großschädl, Stefan Tillich, Alexander Szekely
DSD1
2007 Cryptographic Side-Channels from Low-Power Cache Memory
Philipp Grabher, Johann Großschädl, Dan Page
IMACC2
2007 Instruction Set Extensions for Pairing-Based Cryptography
Tobias Vejda, Dan Page, Johann Großschädl
Pairing3
2007 VLSI Implementation of a Functional Unit to Accelerate ECC and AES on 32-Bit Processors
Stefan Tillich, Johann Großschädl
WAIFI2
2006 Hardware/Software Co-design of Elliptic Curve Cryptography on an 8051 Microcontroller
Manuel Koschuch, Joachim Lechner, Andreas Weitzer, Johann Großschädl, Alexander Szekely, Stefan Tillich, Johannes Wolkerstorfer
CHES4
2006 Instruction Set Extensions for Efficient AES Implementation on 32-bit Processors
Stefan Tillich, Johann Großschädl
CHES2
2006 TinySA: a security architecture for wireless sensor networks
abstract
This paper presents the design and rationale of TinySA, a lightweight security architecture for wireless sensor networks and so-called "smart dust" running the TinyOS operating system. TinySA consists of a suite of security protocols and cryptographic primitives to ensure confidentiality, integrity and authenticity of communication in a sensor network. An integral part of TinySA is a highly optimized elliptic curve cryptosystem, which has been developed from scratch to comply with the extremely limited computational resources available in sensor nodes like the MicaZ mote. This elliptic curve system combines efficient finite field arithmetic with fast curve arithmetic and requires only 5.5.106 clock cycles to compute a 160-bit point multiplication on the Atmega128 processor. Even though our results show that strong elliptic curve cryptography is feasible on sensor nodes, its energy requirements are still orders of magnitude higher compared to that of symmetric cryptosystems. Therefore, TinySA uses elliptic curve cryptography only for infrequent but security-critical operations like key establishment during the initial configuration of the sensor network or the authentication of routing information.
Johann Großschädl
CoNEXT1
2006 Combining algorithm exploration with instruction set design: a case study in elliptic curve cryptography
abstract
In recent years, processor customization has matured to become a trusted way of achieving high performance with limited cost/energy in embedded applications. In particular, instruction set extensions (ISEs) have been proven very effective in many cases. A large body of work exists today on creating tools that can select efficient ISEs given an application source code: ISE automation is crucial for increasing the productivity of design teams. In this paper, we show that an additional motivation for automating the ISE process is to facilitate algorithm exploration: the availability of ISE can have a dramatic impact on the performance of different algorithmic choices to implement identical or equivalent functionality. System designers need fast feedbacks on the ISE-ability of various algorithmic flavors. We use a case study in elliptic curve (EC) cryptography to exemplify the following contributions: (I) ISE can reverse the relative performance of different algorithms for one and the same operation, and (2) automatic ISE, even without predicting speed-ups as precisely as detailed simulation can, is able to show exactly the trends that the designer should follow
Johann Großschädl, Paolo Ienne, Laura Pozzi 0001, Stefan Tillich, Ajay Kumar Verma
DATE1
2005 Energy-Efficient Software Implementation of Long Integer Modular Arithmetic
Johann Großschädl, Roberto Maria Avanzi, Erkay Savas, Stefan Tillich
CHES1
2005 Accelerating AES Using Instruction Set Extensions for Elliptic Curve Cryptography
Stefan Tillich, Johann Großschädl
ICCSA (2)2
2004 Architectural Support for Arithmetic in Optimal Extension Fields
Johann Großschädl, Sandeep S. Kumar, Christof Paar
ASAP1
2004 Instruction Set Extensions for Fast Arithmetic in Finite Fields GF( p) and GF(2m)
Johann Großschädl, Erkay Savas
CHES1
2003 Architectural Enhancements for Montgomery Multiplication on Embedded RISC Processors
Johann Großschädl, Guy-Armand Kamendje
ACNS1
2003 Instruction Set Extension for Fast Elliptic Curve Cryptography over Binary Finite Fields GF(2m)
abstract
The performance of elliptic curve (EC) cryptosystems depends essentially on efficient arithmetic in the underlying finite field. Binary finite fields GF(2/sup m/) have the advantage of "carry-free" addition. Multiplication, on the other hand, is rather costly since polynomial arithmetic is not supported by general-purpose processors. We propose a combined hardware/software approach to overcome this problem. First, we outline that multiplication of binary polynomials can be easily integrated into a multiplier datapath for integers without significant additional hardware. Then, we present new algorithms for multiple-precision arithmetic in GF(2/sup m/) based on the availability of an instruction for single-precision multiplication of binary polynomials. The proposed hardware/software approach is considerably faster than a "conventional" software implementation and well suited for constrained devices like smart cards. Our experimental results show that an enhanced 16 bit RISC processor is able to generate a 191 bit ECDSA signature in less than 650 msec when the core is clocked at 5 MHz.
Johann Großschädl, Guy-Armand Kamendje
ASAP1
2002 Instruction Set Extension for Long Integer Modulo Arithmetic on RISC-Based Smart Cards
abstract
Modulo multiplication of long integers (/spl ges/ 1024 bits) is the major operation of many public-key cryptosystems like RSA or Diffie-Hellman. The efficient implementation of modulo arithmetic is a challenging task, in particular on smart cards due to their constrained resources and relatively slow clock frequency. We present the concept of an application-specific instruction set extension (ISE) for long integer arithmetic. We introduce an optimized multiply-and-accumulate (MAC) unit that makes it possible to compute a/spl times/b+c+d with only one instruction, whereby a, b, c, d are single-precision words (unsigned integers). This additional instruction is simple to incorporate into common RISC architectures like the MIPS32. Experimental results show that the inner-product operation of a multiple-precision multiplication can be accelerated by a factor of two without increasing the processor's clock frequency. We also estimate the execution time of a 1024-bit modulo exponentiation assuming that this special MAC instruction was made available. The proposed ISE is an alternative solution to a crypto co-processor especially for multi-application smart cards (e.g., Java cards) with an embedded 32-bit RISC core.
Johann Großschädl
SBAC-PAD1
2001 A Bit-Serial Unified Multiplier Architecture for Finite Fields GF(p) and GF(2m)
Johann Großschädl
CHES1
2000 The Chinese Remainder Theorem and its Application in a High-Speed RSA Crypto Chip
abstract
The performance of RSA hardware is primarily determined by an efficient implementation of the long-integer modular arithmetic and the ability to utilize the Chinese Remainder Theorem (CRT) for the private key operations. This paper presents the multiplier architecture of the RSA/spl gamma/ crypto-chip, a high-speed hardware accelerator for long-integer modular arithmetic. The RSA/spl gamma/ multiplier datapath is reconfigurable to execute either one 1024-bit modular exponentiation or two 512-bit modular exponentiations in parallel. Another significant characteristic of the multiplier core is its high degree of parallelism. The actual RSA/spl gamma/ prototype contains a 1056/spl times/16-bit word-serial multiplier which is optimized for modular multiplications according to P. Barret's (1987) modular reduction method. The multiplier core is dimensioned for a clock frequency of 200 MHz and requires 227 clock cycles for a single 1024-bit modular multiplication. Pipelining in the highly parallel long-integer unit allows one to achieve a decryption rate of 560 kbit/s for a 1024-bit exponent. In CRT-mode, the multiplier executes two 512-bit modular exponentiations in parallel, which increases the decryption rate by a factor of 3.5 to almost 2 Mbit/s.
Johann Großschädl
ACSAC1
2000 High-Speed RSA Hardware Based on Barret's Modular Reduction Method
Johann Großschädl
CHES1
2000 A New Serial/Parallel Architecture for a Low Power Modular Multiplier
Johann Großschädl
SEC1