Dan Page

dblp:p/DanPage · also Daniel Page · DBLP profile ↗
← Back
43ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0002-6366-7641ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 26 · 2 first-author · 1 since 2021Systems, architecture and hardware · 14 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 1
YearPublicationVenuePosition
2024 RISC-V Instruction Set Extensions for Multi-Precision Integer Arithmetic: A Case Study on Post-Quantum Key Exchange Using CSIDH-512
abstract
Multi-Precision Integer (MPI) arithmetic is a performance-critical component of many public-key cryptosystems, including besides classical ones (e.g., RSA, ECC) also isogeny-based post-quantum schemes. In this paper, we analyze and compare two widely-used MPI representations, namely full-radix and reduced-radix, for the efficient implementation of modular arithmetic operations on the 64-bit RISC-V (RV64GC) architecture. We also evaluate how the execution times of both can be further improved with Instruction Set Extensions (ISEs). The ISEs we propose are able to accelerate a CSIDH-512 class group action by a factor of 1.71 compared to a standard software implementation on a 64-bit Rocket core. This speed-up comes at the cost of a hardware overhead of about 10%.
Hao Cheng 0009, Georgios Fotiadis, Johann Großschädl, Dan Page, Thinh Hung Pham, Peter Y. A. Ryan
DAC4
2022 Towards Micro-architectural Leakage Simulators: Reverse Engineering Micro-architectural Leakage Features Is Practical
Elisabeth Oswald, Dan Page
EUROCRYPT (3)3
2021 A lightweight ISE for ChaCha on RISC-V
abstract
ChaCha is a high-throughput stream cipher designed with the aim of ensuring high-security margins while achieving high performance on software platforms. RISC-V, an emerging, free, and open Instruction Set Architecture (ISA) is being developed with many instruction set extensions (ISE). ISEs are a native concept in RISC-V to support a relatively small RISC-V ISA to suit different use-cases including cryptographic acceleration via either standard or custom ISEs. This paper proposes a lightweight ISE to support ChaCha on RISC-V architectures. This approach targets embedded computing systems such as IoT edge devices that don’t support a vector engine. The proposed ISE is designed to accelerate the computation of the ChaCha block function and align with the RISC-V design principles. We show that our proposed ISEs help to improve the efficiency of the ChaCha block function. The ISE-assisted implementation of ChaCha encryption speeds up at least 5.4× and 3.4× compared to the OpenSSL baseline and ISA-based optimised implementation, respectively. For encrypting short messages, the ISE-assisted implementation gains a comparative performance compared to the implementations using very high area overhead vector extensions.
Ben Marshall, Dan Page, Thinh Hung Pham
ASAP2
2021 XDIVINSA: eXtended DIVersifying INStruction Agent to Mitigate Power Side-Channel Leakage
abstract
Side-channel analysis (SCA) attacks pose a major threat to embedded systems due to their ease of accessibility. Realising SCA resilient cryptographic algorithms on embedded systems under tight intrinsic constraints, such as low area cost, limited computational ability, etc., is extremely challenging and often not possible. We propose a seamless and effective approach to realise a generic countermeasure against SCA attacks. XDIVINSA, an extended diversifying instruction agent, is introduced to realise the countermeasure at the microarchitecture level based on the combining concept of diversified instruction set extension (ISE) and hardware diversification. XDIVINSA is developed as a lightweight co-processor that is tightly coupled with a RISC-V processor. The proposed method can be applied to various algorithms without the need for software developers to undertake substantial design efforts hardening their implementations against SCA. XDIVINSA has been implemented on the SASEBO G-III board which hosts a Kintex-7 XC7K160T FPGA device for SCA mitigation evaluation. Experimental results based on non-specific t-statistic tests show that our solution can achieve leakage mitigation on the power side channel of different cryptographic kernels, i.e., Speck, ChaCha20, AES, and RSA with an acceptable performance overhead compared to existing countermeasures.
Thinh Hung Pham, Ben Marshall, Alexander Fell, Siew-Kei Lam, Dan Page
ASAP5
2015 SoC It to EM: ElectroMagnetic Side-Channel Attacks on a Complex System-on-Chip
Jake Longo, Elke De Mulder, Dan Page, Michael Tunstall
CHES3
2015 Rogue Decryption Failures: Reconciling AE Robustness Notions
Guy Barwell, Dan Page, Martijn Stam
IMACC2
2014 Simulatable Leakage: Analysis, Pitfalls, and New Constructions
Jake Longo, Daniel P. Martin 0001, Elisabeth Oswald, Dan Page, Martijn Stam, Michael Tunstall
ASIACRYPT (1)4
2013 On Secure Embedded Token Design
Simon Hoerder, Kimmo Järvinen 0001, Dan Page
WISTP3
2012 Compiler Assisted Masking
Andrew Moss, Elisabeth Oswald, Dan Page, Michael Tunstall
CHES3
2012 Practical Realisation and Elimination of an ECC-Related Software Bug Attack
Billy Bob Brumley, Manuel Barbosa, Dan Page, Frederik Vercauteren
CT-RSA3
2012 Harnessing Biased Faults in Attacks on ECC-Based Signature Schemes
abstract
This paper presents an extension of the byte-fault attack on signature schemes presented by Giraud et al. Our work extends their attack in a number of ways, but the main focus is an alternative fault model motivated by existing fault injection results. Instead of assuming faults are uniformly distributed (i.e., a given bit is flipped with probability 1/2), we consider the case where faults are biased (i.e., the probability differs from 1/2). Our results show that injecting biased faults allows an attacker to reveal security-critical data with significantly fewer faults and/or a significantly faster search through the remaining candidates.
Kimmo Järvinen 0001, Céline Blondeau, Dan Page, Michael Tunstall
FDTC3
2012 On reconfigurable fabrics and generic side-channel countermeasures
abstract
The use of field programmable devices in security-critical applications is growing in popularity; in part, this can be attributed to their potential for balancing metrics such as efficiency and algorithm agility. However, in common with non-programmable alternatives, physical attack techniques such as fault and power analysis are a threat. We investigate a family of next-generation field programmable devices, specifically those based on the concept of time multiplexing, within this context: our results support the premise that extra, inherent flexibility in such devices can offer a range of possibilities for low-overhead, generic countermeasures against physical attack.
Robert Beat, Philipp Grabher, Dan Page, Stefan Tillich, Marcin Wójcik
FPL3
2012 Efficient Java Implementation of Elliptic Curve Cryptography for J2ME-Enabled Mobile Devices
Johann Großschädl, Dan Page, Stefan Tillich
WISTP2
2011 Bit-Sliced Binary Normal Basis Multiplication
abstract
The performance of many cryptographic primitives is reliant on efficient algorithms and implementation techniques for arithmetic in binary fields. While dedicated hardware support for said arithmetic is an emerging trend, the study of software-only implementation techniques remains important for legacy or non-equipped processors. One such technique is that of software-based bit-slicing. In the context of binary fields, this is an interesting option since there is extensive previous work on bit-oriented designs for arithmetic in hardware, such designs are intuitively well suited to bit-slicing in software. In this paper we harness previous work, using it to investigate bit-sliced, software-only implementation arithmetic for binary fields, over a range of practical field sizes and using a normal basis representation. We apply our results to demonstrate significant performance improvements for a stream cipher, and over the frequently employed Ning-Yin approach to normal basis implementation in software.
Billy Bob Brumley, Dan Page
IEEE Symposium on Computer Arithmetic2
2011 An Exploration of Mechanisms for Dynamic Cryptographic Instruction Set Extension
Philipp Grabher, Johann Großschädl, Simon Hoerder, Kimmo Järvinen 0001, Dan Page, Stefan Tillich, Marcin Wójcik
CHES5
2011 A Unified Multiply/Accumulate Unit for Pairing-Based Cryptography over Prime, Binary and Ternary Fields
abstract
Bilinear maps, or pairings, on elliptic curves are an active area of research in modern cryptology with applications ranging from cryptanalysis (e.g. MOV attack) over identity-based encryption to short signature schemes. Many parameterisations and implementation options for pairing-based cryptography have been investigated in the recent past. Elliptic curves over prime fields are often preferred for software implementation, whereas extension fields of characteristic two and three offer advantages for implementation in hardware. In the ideal case, a hardware accelerator for pairing-based cryptography can support all three types of field to ensure inter-operability with a broad spectrum of applications. This need has motivated the design of so-called unified multipliers, which are basically multipliers that integrate different types of operands (e.g. integers and polynomials) into a single data path. In the present paper, we introduce a unified multiply/accumulate unit for signed/unsigned integers as well as binary and ternary polynomials. The multiplier generates partial products using a Redundant Signed-Digit (RSD) representation that allows for efficient combination of all three operand types into one data path. In addition, our design takes advantage of a high-radix encoding scheme for integers and binary polynomials to reduce the overall number of partial products and utilise the data path in an optimal way. We compare our multiplier with a previous radix-2 implementation of Ozturk et al and analyse the differences in terms of silicon area and critical path delay. The unified multiply/accumulate unit described in this paper can be used in embedded systems like smart cards, either as arithmetic core of a cryptographic co-processor, or as functional unit of an application-specific processor.
Tobias Vejda, Johann Großschädl, Dan Page
DSD3
2011 Can Code Polymorphism Limit Information Leakage?
Antoine Amarilli, Sascha Müller 0003, David Naccache, Dan Page, Pablo Rauzy, Michael Tunstall
WISTP4
2011 An Evaluation of Hash Functions on a Power Analysis Resistant Processor Architecture
Simon Hoerder, Marcin Wójcik, Stefan Tillich, Dan Page
WISTP4
2010 On the Design and Implementation of an Efficient DAA Scheme
Liqun Chen 0002, Dan Page, Nigel P. Smart
CARDIS2
2010 Bridging the gap between symbolic and efficient AES implementations
abstract
The Advanced Encryption Standard (AES) is a symmetric block cipher used to encrypt data within many applications. As a result of its standardisation, and subsequent widespread use, a vast range of published techniques exist for efficient software implementations on diverse platforms. The most efficient of these implementations are written using very low-level approaches; platform dependent assembly language is used to schedule instructions, and most of the cipher is pre-computed into constant look-up tables. The need to resort to such a low-level approach can be interpreted as a failure to provide suitable high-level languages to the cryptographic community. This paper investigates the language features necessary to express AES more naturally (i.e., in a form closer to the original specification) as a source program, and the transformations necessary to produce efficient target programs in an automatic and portable manner.
Andrew Moss, Dan Page
PEPM2
2009 Hardware/Software Co-design of Public-Key Cryptography for SSL Protocol Execution in Embedded Systems
Manuel Koschuch, Johann Großschädl, Dan Page, Philipp Grabher, Matthias Hudler, Michael Krüger
ICICS3
2009 Program interpolation
abstract
Program interpolation is a new type of transformation that given an input program written in a specially constructed Domain Specific Language (DSL), produces a family of functionally equivalent instruction sequences as output. Each sequence is an "interpolation" between the control-flows of implementation strategies supplied in the input program. The purpose of the transformation is to expose behavioural differences (e.g. performance) within the sequences, and thus allow automated optimisation with respect to architectural trade-offs that are difficult to quantify and model. We present results from a prototype compiler that demonstrate a 63% speedup in the domain of multi-precision integer arithmetic.
Andrew Moss, Dan Page
PEPM2
2009 Constructive and Destructive Use of Compilers in Elliptic Curve Cryptography
Manuel Barbosa, Andrew Moss, Dan Page
J. Cryptol.3
2008 Light-Weight Instruction Set Extensions for Bit-Sliced Cryptography
Philipp Grabher, Johann Großschädl, Dan Page
CHES3
2008 Randomised representations
abstract
The authors show that a number of existing methods for side-channel defence are essentially the same techniques presented in different contexts. By abstracting this technique, they present necessary conditions which need to be satisfied for it to be successful in preventing side-channel analysis. They also show that concrete application of the technique via randomised field representation produces more efficient implementations than application of the technique via randomised projective coordinates.
Nigel P. Smart, Elisabeth Oswald, Dan Page
IET Inf. Secur.3
2007 Cryptographic Side-Channels from Low-Power Cache Memory
Philipp Grabher, Johann Großschädl, Dan Page
IMACC3
2007 Toward Acceleration of RSA Using 3D Graphics Hardware
Andrew Moss, Dan Page, Nigel P. Smart
IMACC2
2007 Instruction Set Extensions for Pairing-Based Cryptography
Tobias Vejda, Dan Page, Johann Großschädl
Pairing2
2007 Nondeterministic Multithreading
abstract
The physical security of application-specific embedded processors such as those found in smart-cards has become increasingly important since they are used more and more as conduits for sensitive financial and identity information. The advent of side-channel attacks has meant that a combination of algorithm, software, and hardware defense is required. In this paper, we reexamine the issue of nondeterministic processors, simplifying previous designs using a multithreaded architecture. From this simplification, we are able to construct a formally reasoned assessment of the security level offered by such a device.
Peter James Leadbitter, Dan Page, Nigel P. Smart
IEEE Trans. Computers2
2006 DFM: swimming upstream
abstract
In this tutorial we will examine the process issues that are causing yield stability issues and how they will affect design flows for 65nm and below.
Dan Page, Jamil Kawa, Charles C. Chiang
ACM Great Lakes Symposium on VLSI1
2006 A Fault Attack on Pairing-Based Cryptography
abstract
Current fault attacks against public key cryptography focus on traditional schemes, such as RSA and ECC, and, to a lesser extent, on primitives such as XTR. However, bilinear maps, or pairings, have presented theorists with a new and increasingly popular way of constructing cryptographic protocols. Most notably, this has resulted in efficient methods for identity based encryption (IBE). Since identity-based cryptography seems an ideal partner for identity aware devices such as smart-cards, in this paper, we examine the security of concrete pairing instantiations in terms of fault attack
Dan Page, Frederik Vercauteren
IEEE Trans. Computers1
2005 Hardware Acceleration of the Tate Pairing in Characteristic Three
Philipp Grabher, Dan Page
CHES2
2005 Practical Cryptography in High Dimensional Tori
Marten van Dijk, Robert Granger, Dan Page, Karl Rubin, Alice Silverberg, Martijn Stam, David P. Woodruff
EUROCRYPT3
2005 On the Automatic Construction of Indistinguishable Operations
Manuel Barbosa, Dan Page
IMACC2
2005 Hardware and Software Normal Basis Arithmetic for Pairing-Based Cryptography in Characteristic Three
abstract
Although identity-based cryptography offers a number of functional advantages over conventional public key methods, the computational costs are significantly greater. The dominant part of this cost is the Tate pairing, which, in characteristic three, is best computed using the algorithm of Duursma and Lee. However, in hardware and constrained environments, this algorithm is unattractive since it requires online computation of cube roots or enough storage space to precompute required results. We examine the use of normal basis arithmetic in characteristic three in an attempt to get the best of both worlds: an efficient method for computing the Tate pairing that requires no precomputation and that may also be implemented in hardware to accelerate devices such as smart-cards.
Robert Granger, Dan Page, Martijn Stam
IEEE Trans. Computers2
2004 Attacking DSA Under a Repeated Bits Assumption
Peter James Leadbitter, Dan Page, Nigel P. Smart
CHES2
2004 Parallel Cryptographic Arithmetic Using a Redundant Montgomery Representation
abstract
We describe how using a redundant Montgomery representation allows for high-performance SIMD-based implementations of RSA and elliptic curve cryptography. This is in addition to the known benefits of immunity from timing attacks afforded by the use of such a representation. We present some preliminary implementation timings using the SSE2 instruction set on a Pentium 4 processor and show that an SIMD parallel implementation of RSA can be around twice as fast as traditional sequential code. This is especially useful given the larger 2,048 bit RSA keys which are now being proposed for standard security levels. Finally, we remark on other application areas that improve the security of our work in the context of side-channel analysis while maintaining high performance.
Dan Page, Nigel P. Smart
IEEE Trans. Computers1
2003 Using Media Processors for Low-Memory AES Implementation
abstract
Most performance studies of AES make traditional space versus time tradeoffs by allowing large lookup tables to accelerate operations that would normally be calculated by the processor. However, AES is a versatile algorithm and can also be optimised for low-memory use in constrained environments. We investigate the possibility of getting the best of both worlds - an application specific hardware and software solution that has a low dependency on memory yet still executes fast enough to consider for use in production systems. The resulting software is attractive in high level design since it allows AES to be more easily deployed as a composable element in larger systems and scale better as processor speed increases.
James Irwin, Dan Page
ASAP2
2003 Defending against cache-based side-channel attacks
Dan Page
Inf. Secur. Tech. Rep.1
2002 Predictable Instruction Caching for Media Processors
abstract
The determinism of instruction cache performance can be considered a major problem in multimedia devices which hope to maximise their quality of service. If instructions are evicted from the cache by competing blocks of code, the running application will take significantly longer to execute than if the instructions were present. Since it is difficult to predict when this interference will occur the performance of the algorithm at a given point in time is unclear We propose the use of an automatically configured partitioned cache to protect regions of the application code from each other and hence minimise interference. As well as being specialised to the purpose of providing predictable performance, this cache can be specialised to the application being run, rather than for the average case, using simple compiler algorithms.
James Irwin, David May 0001, Henk L. Muller, Dan Page
ASAP4
2002 Instruction Stream Mutation for Non-Deterministic Processors
abstract
Differential power analysis (DPA) has become a real-world threat to the security of cryptographic hardware devices such as smart-cards. By using cheap and readily available equipment, attacks can easily compromise algorithms running on these devices in a non-invasive manner. Adding non-determinism to the execution of cryptographic algorithms has been proposed as a defence against these attacks. One way of achieving this non-determinism is to introduce random additional operations to the algorithm which produce noise in the power profile of the device. We describe the addition of a specialised processor pipeline stage which increases the level of potential non-determinism and hence guards against the revelation of secret information.
James Irwin, Dan Page, Nigel P. Smart
ASAP2
2002 Hardware Implementation of Finite Fields of Characteristic Three
Dan Page, Nigel P. Smart
CHES1
1999 Microcaches
David May 0001, Dan Page, James Irwin, Henk L. Muller
HiPC2