EDBT 2026 Demo / reviewers in the wild / expert
Máire O'Neill
dblp:m/MaireMcLoone · also Máire McLoone
· DBLP profile ↗
105ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0002-6865-6212ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 85 · 6 first-author · 27 since 2021Security and privacy · 12 · 2 first-author · 4 since 2021Computer networks · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Planting CRYSTALS-Kyber Acceleration SeedsabstractCRYSTALS-Kyber (FIPS203) was one of the first post quantum cryptography schemes to be standardized by the National Institute of Standards and Technology; CRYSTALS-Kyber utilizes Montgomery reduction at the heart of the modulo reduction. Montgomery reduction traditionally is suited to computing moduli much larger than 32/64-bits. Recently a new modulo reduction method was published named Plantard reduction which allows for one less multiplication at the reduction stage. However, this does require doubling the word size of the values being calculated, in the case of CRYSTALS-Kyber increasing from 16-bit to 32-bit. In this paper we present the worlds first Plantard ISE on the Ibex RISC-V core and investigate the impact Instruction Set Extensions can have utilizing this new modulo reduction method over original methods such as Montgomery and Barrett reduction. Utilizing Plantard reduction as a set of Instruction Set Extensions, we are able to reduce the cycle counts required to compute the entire CRYSTALS-Kyber algorithm by up to 21% and reduce the cycle counts of the NTT by 1.7x and the INTT by 2.6x against the reference implementation. Ryan Bevin, Ayesha Khalid, Seog Chung Seo, Máire O'Neill |
ISCAS | 4 |
| 2026 | GUARDIAN: A Decoupled GNN Framework for Type-aware Hardware Trojan DetectionabstractWith increased outsourcing of intellectual property (IP) cores to third-party vendors, system-on-chip designs face heightened security risks from hardware trojans (HTs) that can compromise integrity and functionality. Prior HT detection works typically train separate deep-learning models on combinational (TRTC-TC) and sequential (TRTC-TS) trojan benchmarks from the LEDA-based Trust-Hub dataset, treating them independently and framing detection as binary classification, which fails when both trojan types coexist in a single netlist. In contrast, our proposed method employs a unified Graph Neural Network (GNN)-based model trained on a mixed dataset comprising benchmarks from both TRTC-TC and TRTC-TS, enabling multi-class classification to distinguish between benign, combinational, and sequential trojans. The approach addresses key GNN constraints and attains strong detection performance on the mixed dataset: TPR of 99.3% and 100%, TNR of 99.99% and 99.9%, and PPV of 97.9% and 97% for combinational and sequential trojans, respectively. Zain Shabbir, Shichao Yu, Ihsen Alouani, Máire O'Neill |
ISCAS | 4 |
| 2026 | HBQS: Lightweight Post-Quantum Secure Authentication for Satellite Networks Leveraging Hardware TRNG and PUFsabstractSatellite communication networks play a critical role in providing connectivity to remote regions and areas with limited infrastructure. However, their inherently open nature and physical exposure make them particularly susceptible to security threats, including replay, impersonation, and man-in- the-middle attacks. The emergence of quantum computing further undermines the robustness of conventional cryptographic schemes that rely on number-theoretic assumptions. To mitigate these challenges, this article proposesHBQS, a lightweight post-quantum authentication framework designed for satellite platforms with limited resources.HBQSintegrates Physically Unclonable Functions (PUFs) with hash-based cryptography, leveraging SPHINCS+ digital signatures and SHA-3 hashing to provide secure mutual authentication and session key establishment. The protocolHBQSwas implemented and evaluated on embedded hardware platforms, including Raspberry Pi 4.0, PYNQ-Z2 FPGA, and a Dell ground control station. The experimental results demonstrate thatHBQSachieves mutual authentication in 0.88 ms, offering an approximately 67% reduction in execution time compared to representative baselines from the prior literature. Entropy analysis confirms that the proposed protocol maintains a high entropy across critical components, withHBQSachieving a signature entropy of 164.63 bits and PUF response entropy of 172.58 bits, indicating strong resistance to statistical and modeling attacks. A formal security analysis conducted within the Random Oracle Model demonstrates semantic security against both classical and quantum adversaries. The protocolHBQSshows a significant improvement over existing methods, achieving approximately 67% reduction in computational execution time, approximately 11.5% faster user-side handshake time, approximately 0.6% improvement on the satellite side and full security coverage across all evaluated security features with only a 21.4% increase in static memory usage as the trade-off. These findings positionHBQSas an efficient, secure, and scalable authentication solution for next-generation satellite communication systems operating in the post-quantum era. Muhammad Arslan Akram, Arnab Kumar Biswas, Máire O'Neill, Ayesha Khalid, Adnan Noor Mian |
IEEE Internet Things J. | 3 |
| 2026 | LightHD: A Lightweight and High-Performance Hardware Accelerator of CRYSTALS-DilithiumabstractCRYSTALS-Dilithium serves as the foundation of the NIST-standardised PQC digital signature scheme, and has been declared as the first recommended digital signature algorithm. However, due to the computational complexity and intricate processing flow of CRYSTALS-Dilithium, two limitations are shown in existing methods: its applicability on resource-constrained devices is limited and the performance reported so far remains relatively low. This paper presents a lightweight yet high-performance hardware architecture that optimises the core computational units of CRYSTALS-Dilithium. First, an iterative dual-Keccak SHA-3 module is proposed, where two cores operate with a 26-cycle offset to accelerate processing without compromising frequency. In addition, the rejection sampler is streamlined by two compact registers for intermediate values and counters, improving efficiency when consuming interleaved SHA-3 outputs. Second, for small bit-width polynomials, we eliminate the first NTT stage via lookup tables and data regrouping, reducing NTT cycles by 11.7% with little hardware overhead. Further hardware savings are achieved by maximising IP core utilisation and simplifying input multiplexers. Furthermore, a compact scheduling strategy ensures that all intermediate storage fits within a single polynomial-sized memory block. On Xilinx Artix-7 FPGAs, the design reduces hardware overhead by 14.2% compared with state-of-the-art lightweight implementations. Across three security levels, KeyGen and Verify are 27.3% and 13.5% faster, respectively, than high-performance prior designs. At level 5, the best-case Sign latency is only 120 µs. Ziying Ni, Ayesha Khalid, Zhaoyu Zhang 0001, Yijun Cui, Weiqiang Liu 0001, Máire O'Neill |
IEEE Trans. Computers | 6 |
| 2026 | A Methodology for Pre-Silicon Optimization of Processor Based PUF in Approximate ComputingabstractThe unpredictable inherent error behavior of approximate computing introduces both new security threats and opportunities to design novel security primitives/strategies. This work proposes a methodology that exploits stochastic timing errors of a pipelined datapath caused by voltage scaling to design an optimized processor-based physical unclonable function (PUF) for approximate computing. To verify the effectiveness of this method, a pipelined arithmetic architecture is implemented at a 45 nm technology node, and voltage scaling is applied to extract PUF bits. With reduced supply voltage, harvested PUF bits show increased uniqueness. Moreover, proposed divergent delay path selection based on intermediary error behavior exhibits improved PUF uniqueness vs an unmodified datapath. A design optimization methodology is applied introducing new PUF metrics - gain (G) and performance power ratio (PPR). Using these metrics, the optimum scaled voltage range is identified for enhanced PUF performance. The optimized PUF shows maximum uniqueness of 49%, and reliability of 92% with a temperature range of -20${\circ }$C to 70${\circ }$C. Further, the proposed PUF with approximate computing achieves markedly improved G and PPR relative to the exact case. With better uniqueness, reliability, and low resource utilization, the proposed PUF methodology is highly suitable for securing approximate computing applications. Aditya Japa, Robert James Moore, Jack Miskelly, Jiliang Zhang 0002, Weiqiang Liu 0001, Máire O'Neill, Chongyan Gu |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | DeepPUFSCA: Deep learning for Physical Unclonable Function attack based on Side Channel Analysis supportabstractPhysical Unclonable Function (PUF) poses a vulnerability that it could be imitated by machine learning attacks and side channel attacks, which break its physical uniqueness and unpredictable characteristic. Hence, many works are concerned with enhancing PUF design by introducing more nonlinear modules inside to differentiate approximating PUF behavior from the attacker side. However, the safety of these PUFs are still an open area and need to be verified. In this paper, we propose DeepPUFSCA, which is a deep learning-based model that uniquely combines both challenge and side-channel information features during training to attack PUF. To gather the data, we conduct a design of an arbiter PUF on FPGA and measure its power consumption. Our intensive experiments on this dataset demonstrate that DeepPUFSCA outperforms other machine learning-based methods in terms of attacking accuracy, even the novel ensemble algorithms. Moreover, we also show that combined side channel information boosts the model performance compared to attacking with challenge-response only. Ngoc Phu Doan, Tuan Dung Pham, Zichi Zhang, Viet-Hung Tran, Jack Miskelly, Hans Vandierendonck, Anh-Tuan Hoang, Máire O'Neill, Son T. Mai |
DAC | 8 |
| 2025 | Security of Approximate Neural Networks against Power Side-channel AttacksabstractEmerging low-energy computing technologies, in particular approximate computing, are becoming increasingly relevant in key applications. A significant use case for these technologies is reduced energy consumption in Artificial Neural Networks (ANNs), an increasingly pressing concern with the rapid growth of AI deployments. It is essential we understand the security implications of approximate computing in an ANN context before this practice becomes commonplace. In this work, we examine the test case of approximate ANN processing elements (PE) in terms of information leakage via the power side channel. We perform a weight extraction correlation Power Analysis (CPA) attack under three approximation scenarios: overclocking, voltage scaling, and circuit level bitwise approximation. We demonstrate that as the degree of approximation increases the Signal to Noise Ratio (SNR) of power traces rapidly degrades. We show that the Measurement to Disclosure (MTD) increases for all approximate techniques. An MTD of 48 under precise computing is increased to at minimum 200 (bitwise approximate circuit at $\mathbf{2 5 \%}$ approximation), and under some approximation scenarios $\gt1024$. i.e. an increase in attack difficulty of at least x4 and potentially over x20. A relative Security-Power-Delay (SPD) analysis reveals that, in addition to the across the board improvement vs precise computing, voltage and clock scaling both significantly outperform approximate circuits with voltage scaling as the highest performing technique. Aditya Japa, Jack Miskelly, Máire O'Neill, Chongyan Gu |
DAC | 3 |
| 2025 | Invited Paper: Rowhammer Mitigation by Approximate Computing: A Compressed Sensing Case StudyabstractWhile Approximate Computing (AC) trades the precision for energy efficiency with tolerable errors, its security of approximate data in digital storage is not well explored for edge devices. As one of the most effective hardware security attack methods, Rowhammer attack has shown significant threats to the digital data on dynamic random access memory (DRAM) with the vulnerability of high-frequency memory row access. This work performs the first preliminary evaluation of Rowhammer attack on real-world compressed sensing applications with approximate data. By investigating Rowhammer attack on the approximate data from compact LiDar sensor, the security impact of various precisions is presented through the fidelity of reconstructed depth image. The experiments reveal considerable mitigation of Rowhammer attack by adopting AC based sensor signal processing, where up to 2× higher PSNR of output depth image is achieved comparing to those with accurate data and computations. Yuhang Hao, Yun Wu 0003, Minmin Jiang, Máire O'Neill, Chongyan Gu |
ICCAD | 4 |
| 2025 | AxRA: Approximate Rowhammer Attack for Modern DRAM SystemsabstractApproximate computing achieves high performance or less power consumption in various fault-tolerant applications, e.g., image processing, artificial intelligence (AI), etc. However, the introduction of approximate computing brings new security vulnerabilities, which threaten the entire computing system. In this paper, a novel Rowhammer attack is proposed, which utilises the approximate data stored in DRAM memories to achieve higher attack effectiveness. Compared to Rowhammer attack to DRAM memory without approximate data, the proposed method achieves more bit-flips resulting in significant data corruption. The proposed attack is implemented and evaluated on DRAM chips with a real user case, object detection using neural network. The accuracy of detection on the baseline image is employed to verify the impact of proposed attack approach. The results show that the proposed Rowhammer attack with approximate data introduces extra 33% bit-flips on victim rows than a conventional Rowhammer attack without approximate data. It also introduces up to ∼75% accuracy reduction of MNIST neural network proportionally to the increment of attack activation number. Yuhang Hao, Yun Wu 0003, Ziying Ni, Jack Miskelly, Máire O'Neill, Chongyan Gu |
ISCAS | 5 |
| 2025 | Kyber-KEM-Ascon: Benchmarking a Lightweight Post-quantum KEM on IoT DevicesabstractCRYSTALS-Kyber, officially standardized by the U.S. National Institute of Standards and Technology (NIST) in August 2024 as ML-KEM FIPS, is the only post quantum secure key encapsulation mechanism (KEM) standard. Due to its inherent computationally intensive nature, mapping it on lightweight IoT devices is often a struggle. This study examines the replacement of the Keccak hashing function in Kyber-KEM with the NIST LWC competition winner called Ascon hash function. We compare speed/ memory improvement by executing Kyber-KEM-Ascon on an ARM Cortex M4 device and present novel benchmarking results; Kyber-KEM-Ascon shows about 25% reduction in clock cycles, along with lower memory usage requirements (about 8%). These results suggest that Kyber-KEM-Ascon is more suited for lightweight platforms, offering benefits for IoT security applications. Sahar Shehzadi, Nathan Whaley, Ayesha Khalid, Abdul Ghafoor 0002, Sadiqa Arshad, Faiz Ul Islam, Máire O'Neill |
ISCAS | 7 |
| 2025 | An Enhanced Two-Step CPA Side-Channel Analysis Attack on ML-KEMabstractThis work presents an enhanced two-step Correlation Power Analysis (CPA) attack targeting the recently standardised ML-KEM on an ARM Cortex M4. Our enhancement exploits the knowledge of intermittent variables to identify sample points of interest and develop bespoke attack functions. Step one targets the odd coefficients of each Secret Key Polynomial Vector ( ˆ s), before step two targets the remaining even coefficients using more elaborate attack functions. After successfully demonstrating key recovery for the first set of ˆ s, we then characterise leakage behaviour, revealing a trend indicating recovery of each coefficient becomes more efficient with subsequent iterations of the internal doublebasemul operation. By applying our enhanced two step attack methodology, we successfully recovered the entire key using only 179 traces, without the need for elaborate preconditions or ciphertext manipulations. We obtain remarkable results in the initial stage of our attack, while the second phase achieves performance comparable to other recent studies. Mark Kennaway, Anh-Tuan Hoang, Ayesha Khalid, Ciara Rafferty, Máire O'Neill |
SECRYPT | 5 |
| 2025 | A Highly Hardware Efficient ML-KEM Accelerator with Optimised Architectural LayersabstractThe Module-Lattice-Based Key encapsulation Mechanism (ML-KEM) scheme, which is currently being standardised, is a quantum attack resistant KEM that is based on CRYSTALS-Kyber. CRYSTALS-Kyber is the only Public-key Encryption (PKE)/ KEM scheme selected in the first set of successful candidates as part of the NIST initiated Post-Quantum Cryptography (PQC) process. ML-KEM scheme includes three different security levels, namely security level 1, 3, and 5. In this research, we propose a highly area-time efficient hardware ML-KEM architecture. The architecture comprises three computational layers. The first layer comprises a hash and sampling module; the second layer includes a number theoretic transform (NTT), its inverse (INTT) and a point-wise multiplication (PWM) module; and the third layer comprises addition, compressing and encoding. Intra-layer pipelining and out-of-layer scheduling ensures that either layer 1 or layer 2 operate in the shortest time. In the reduction module, we propose a novel hybrid architecture to obtain the final result within 2 cycles with low area consumption. In the NTT module, the PWM pipelining method is modified and an optimised iterative FIFO access method is adopted to reduce the size of FIFO units by 55% over previous research. Look-up tables are also used to replace the first-stage of the NTT to reduce 8 cycles. Furthermore, the memory unit uses only FIFOs, the size are optimised based on the requirements of the most resource-intensive function in ML-KEM (ML-KEM.CPA.Dec). The results show that the proposed architecture has a 48.2%, 41.2%, and 78.1% reduction in computational time in comparison to previous work for security levels 1, 3, and 5, respectively. In addition, the area of proposed optimised ML-KEM designs is reduced by 73%, 70%, 76% and resulting in an improved area-time (AT) product of 15.8%, 10.7%, and 11.3%, for the Level 1, 3, and 5 security levels respectively, compared with state-of-the-art designs. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2024 | Bitstream Fault Injection Attacks on CRYSTALS Kyber Implementations on FPGAsabstractCRYSTALS-Kyber is the only Public-key Encryption (PKE)/ Key-encapsulation Mechanism (KEM) scheme that was chosen for standardization by the National Institute of Standards and Technology initiated Post-quantum Cryptography competition (so called NIST PQC). In this paper, we show the first successfully malicious modifications of the bitstream of a Kyber FPGA implementation. We successfully demonstrate 4 different attacks on Kyber hardware implementations on Artix-7 FPGAs that either reduce the complexity of polynomial multiplication operations or enable direct secret key/ message recovery by: disabling BRAMs, disabling DSPs, zeroing NTT ROM and tampering with CBD2 results. Two of our attacks are generic in nature and the other two require reverse-engineering or a detailed knowledge of the design. We evaluate the feasibility of the four attacks, among which the zeroing NTT ROM and tampering with the CBD2 result attacks produce higher public key and ciphertext complexity and thus are difficult to be detected. Two countermeasures are proposed to prevent the attacks proposed in this paper. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
DATE | 4 |
| 2024 | A Novel Methodology for Processor based PUF in Approximate ComputingabstractApproximate computing has great potential in the design of high-performance and energy-efficient systems. The inherent stochastic error behavior of approximate computing introduces both new security threats and opportunities to enhance security. This work proposes a novel methodology that exploits stochastic timing errors of a pipelined datapath to design a processor based physically unclonable function (PUF) for approximate computing. This methodology uses divergent delay path selection based on intermediary error behaviour to improve the PUF uniqueness vs. an unmodified datapath, even when only moderate voltage scaling is applied. To verify the effectiveness of this method, a pipelined fast fourier transform (FFT) butterfly architecture is implemented at 45nm technology node, and a voltage over scaling technique is applied to extract PUF bits. The proposed methodology achieves a maximum uniqueness of 48.5% whereas conventional design uniqueness is limited to 43%. Overall, the proposed design shows a maximum of ~7% higher uniqueness and ~10% higher reliability (for iso uniqueness) compared to the conventional pipelined design. Aditya Japa, Jack Miskelly, Yijun Cui, Máire O'Neill, Chongyan Gu |
ISCAS | 4 |
| 2024 | FPGA Bitstream Fault Injection Attack and Countermeasures on the Sampling Counter in CRYSTALS KyberabstractThe CRYSTALS Kyber algorithm is the public key encryption (PKE)/ key encapsulation mechanism (KEM) protocol undertaken for standardization by the US National Institute of Standards and Technology (NIST) the PQC competition and serves as the foundation for the Module-Lattice-Based (ML)-KEM scheme. The inherently strong security properties of the Kyber algorithm are considered to be resistant to attacks under quantum computers, but the security of its FPGA-based hardware implementation circuitry is still worth considering. In this work, we introduce the Nonce counter disabling attack, which targets the binomial distribution sampling process. We demonstrate that, in the modified primes of Kyber from Round 2, it is also effectively deduce the secret key s by equating it with the noise e. Our implementation of this attack on a Nexys 4 FPGA, with an additional DSP disabling filtering process to pinpoint the LUT. This attack is applicable to both the key generation and key encapsulation phases, and only need to modify 32-bit bitstream. Finally, We propose the Nonce counter check and the splitting of the Nonce computation cycles methods to to prevent this attack in hardware design-level. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 4 |
| 2024 | Efficient Soft Core Multiplier for Post Quantum Digital SignaturesabstractMultiplication is a core operation in various applications such as cryptography and machine learning. Dedicated DSP blocks are provided by FPGA vendors for multiplication. However, these DSP blocks are limited in number and their location on FPGA is fixed, resulting in routing delays that affects the performance for small size multipliers. In this paper, a high performance and resource efficient 5 × 5 multiplier is presented that utilizes lookup tables (LUTs) and fast carry chain of the FPGA. The proposed multiplier offers 30% reduction in LUTs compared to Vivado DSP-less inferred multiplier at the cost of a slight increase in critical path delay (CPD). The proposed multiplier requires lesser power consumption and has better area- delay product (ADP) and power-delay product (PDP) metrics. Based on the proposed multiplier, a finite field multiplier is developed for post quantum digital signatures such as QR-UOV, MAYO and MQOM. The matrix-vector architecture is the core operation in multivariate digital signatures and integration of our finite field multiplier in a matrix-vector architecture shows that area is almost halved compared to state-of-the-art. Yasir Ali Shah, Ciara Rafferty, Ayesha Khalid, Safiullah Khan, Khalid Javeed, Máire O'Neill |
ISCAS | 6 |
| 2024 | Quantum-Safe HIBE: Does It Cost a Latte?abstractThe United Kingdom (UK) government is considering advanced primitives such as identity-based encryption (IBE) for adoption as they transition their public-safety communications network from TETRA to an LTE-based service. However, the current LTE standard relies on elliptic-curve-based IBE, which will be vulnerable to quantum computing attacks, expected within the next 20–30 years. Lattices can provide quantum-safe alternatives for IBE. These schemes have shown promising results in terms of practicality. To date, several IBE schemes over lattices have been proposed, but there has been little in the way of practical evaluation. This paper provides the first complete optimised practical implementation and benchmarking of Latte, a promising Hierarchical IBE (HIBE) scheme proposed by the UK National Cyber Security Centre (NCSC) in 2017 and endorsed by European Telecommunications Standards Institute (ETSI). We propose optimisations for the KeyGen, Delegate, Extract and Gaussian sampling components of Latte, to increase attack costs, reduce decryption key lengths by 2x–3x, ciphertext sizes by up to 33%, and improve speed. In addition, we conduct a precision analysis, bounding the Rényi divergence of the distribution of the real Gaussian sampling procedures from the ideal distribution in corroboration of our claimed security levels. Our resulting implementation of the Delegate function takes 0.4 seconds at 80-bit security level on a desktop machine at 4.2GHz, significantly faster than the order of minutes estimated in the ETSI technical report. Furthermore, our optimised Latte Encrypt/Decrypt implementation reaches speeds up to 9.7x faster than the ETSI implementation. Raymond K. Zhao, Sarah McCarthy, Ron Steinfeld, Amin Sakzad, Máire O'Neill |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | An Efficient Ring Oscillator PUF Using Programmable Delay Units on FPGAabstractThe ring oscillator (RO) PUF can be implemented on different FPGA platforms with high uniqueness and reliability. To decrease the hardware cost of conventional RO PUFs, a new design using the programmable delay units is proposed, namely, PRO PUF. The programmable interconnect points (PIPs) of programmable delay units are used to enhance the configurability. The PUF cell of the proposed design has the ability to be efficiently programmed to an RO PUF at any stage by adjusting the propagation paths of the delay units. A significant number of responses can be generated by the proposed PRO PUF while consuming fewer hardware resources. To verify the performance, the proposed design has been implemented on Xilinx FPGAs and also simulated using a standard 40nm technology. The experimental results have shown that the proposed design achieves high uniqueness, reliability, and hardware efficiency. Moreover, the PRO PUF has been evaluated using a machine learning attack, the CMA-ES attack. The results have shown that the proposed structure is more resistant to common modeling attacks when compared to conventional RO-related PUF designs. Yijun Cui, Jiang Li 0012, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2023 | Towards a Lightweight CRYSTALS-Kyber in FPGAs: an Ultra-lightweight BRAM-free NTT CoreabstractCRYSTALS-Kyber is the first quantum-resilient, lattice-based Public Key Encryption (PKE)/Key Encapsulation Mechanism (KEM) cryptosystem that is chosen by the ongoing National Institute of Standards and Technology post-quantum cryptography standardization (NIST PQC) for standardization. This work presents a lightweight and efficient, FPGA-based hardware implementation for polynomial multiplication unit (NTT), which is the major bottleneck in the Kyber scheme. As a first step, an optimzed modular multiplication architecture combining KRED and lookup table-based algorithms is presented, which reduces the resources of slices by 16.7%. It is used in a pipelined NTT/INTT architecture that is completely BRAM free and instead uses 3 FIFOs for coefficients storage. We hereby present the most compact FPGA based design for NTT architecture in Kyber till date. Experimental results bench marked on comparable FPGA devices show that our proposed design is 36-75% better than the state-of-the-art implementations in terms of hardware efficiency for NTT/INTT calculations and$3.4-4.4\times$better for the Point-wise Multiplication (PWM) operation. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 4 |
| 2023 | Efficient, Error-Resistant NTT Architectures for CRYSTALS-Kyber FPGA AcceleratorsabstractThe dawn of cost-effective miniaturised satellites is currently attracting venture capital in a never seen before ratio to launch mega-constellations of satellites for a diverse range of applications. These satellites are vulnerable to attacks by high-capability cyber-criminals (including quantum enabled adversaries), due to the critical data they transmit. Additionally, space missions have long lifespan and a long lead time in terms of development process, requiring a pre-emptive outlook to ensuring their safety. In 2016, National Institute of Standards and Technology (NIST) initiated the competition to standardise the post-quantum cryptography (PQC) schemes, announcing the first portfolio of chosen schemes in 2022. This work targets the only public key exchange (PKE) scheme among the winners of the NIST-PQC standardisation process, CRYSTALS-Kyber, and implements its core bottleneck operation, i.e., number theoretic transform (NTT) extensively used for the polynomial multiplication. To avoid data corruption due to space based radiations, a novel error-resistant model for NTT is presented based on hybrid protection mechanisms, i.e., the use of hamming codes for detection and correction of errors in the twiddle factors and the use of parity computed for all NTT coefficients for error detection. Benchmarking error protection overheads on a Xilinx Virtex-7 FPGA reports 16.4% and 10.8% degradation on the hardware efficiency when the hamming codes for twiddle factors and parity bit for NTT coefficients are used to mitigate errors, respectively. A total of 29.2% area overhead is benchmarked when compared to the standard unprotected NTT implementations. Safiullah Khan, Ayesha Khalid, Ciara Rafferty, Yasir Ali Shah, Máire O'Neill, Wai-Kong Lee, Seong Oun Hwang |
VLSI-SoC | 5 |
| 2023 | HPKA: A High-Performance CRYSTALS-Kyber Accelerator Exploring Efficient PipeliningabstractCRYSTALS-Kyber (Kyber) was recently chosen as the first quantum resistant Key Encapsulation Mechanism (KEM) scheme for standardisation, after three rounds of the National Institute of Standards and Technology (NIST) initiated PQC competition which begin in 2016 and search of the best quantum resistant KEMs and digital signatures. Kyber is based on the Module-Learning with Errors (M-LWE) class of Lattice-based Cryptography, that is known to manifest efficiently on FPGAs. This work explores several architectural optimizations and proposes a high-performance and area-time (AT) product efficient hardware accelerator for Kyber. The proposed architectural optimizations include inter-module and intra-module pipelining, that are designed and balanced via FIFO based buffering to ensure maximum parallelisation. The implementation results show that compared to state-of-the-art designs, the proposed architecture delivers 25–51% speedups for Kyber's three different security levels on Artix-7 and Zynq UltraScale+ devices, and a 50–75% reduction in DSPs at comparable security level. Consequently, the proposed design achieve higher AT product efficiencies of 19–33%. Ziying Ni, Ayesha Khalid, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Computers | 4 |
| 2022 | Acceleration of Post Quantum Digital Signature Scheme CRYSTALS-Dilithium on Reconfigurable HardwareabstractThis research investigates efficient architectures for the implementation of the CRYSTALS-Dilithium post-quantum digital signature scheme on reconfigurable hardware, in terms of speed, memory usage, power consumption and resource utilisation. Post quantum digital signature schemes involve a significant computational effort, making efficient hardware accelerators an important contributor to future adoption of schemes. This is work in progress, comprising the establishment of a comprehensive test environment for operational profiling, and the investigation of the use of novel architectures to achieve optimal performance. Donal Campbell, Ciara Rafferty, Ayesha Khalid, Máire O'Neill |
FPL | 4 |
| 2022 | High Performance FPGA-based Post Quantum Cryptography ImplementationsabstractPost-quantum Cryptography (PQC) is an umbrella term for cryptographic schemes based on hard mathematical problems which are resistant to attacks by quantum computers. The National Institute of Standards and Technology (NIST) initiated a PQC standardisation process in 2017, with a total of 4 algorithms selected for standardisation after round 3 and 4 undertaken for further analysis in Round 4 in 2022. PQC schemes on hardware devices, such as Field Programmable Gate Arrays (FPGA), show the potential of higher throughput performance, for comparable security, at the cost of high area and power consumption. The major aim of this thesis is to help facilitate the global transition to a post quantum secure set of security protocols. This thesis will focus on the optimisation of the the hardware architectures to improve the computational speed and reduce the area overhead. The side channel analysis vulnerabilities and their countermeasures will also be studied. Ziying Ni, Ayesha Khalid, Máire O'Neill |
FPL | 3 |
| 2022 | Stacked Ensemble Model for Enhancing the DL based SCAabstractDeep learning (DL) has proven to be very effective for image recognition tasks, with a large body of research on various models for object classification. The application of DL to side-channel analysis (SCA) has already shown promising results, with experimentation on open-source variable key datasets showing that secret keys for block ciphers like Advanced Encryption Standard (AES)-128 can be revealed with 40 traces even in the presence of countermeasures. This paper aims to further improve the application of DL in SCA, by enhancing the power of DL when targeting the secret key of cryptographic algorithms when protected with SCA countermeasures. We propose a stacked ensemble model, which trains the output probabilities and Maximum likelihood score of multiple traces and/or sub-models to improve the performance of Convolutional Neural Network (CNN)-based models. Our model generates state-of-the art results when attacking the ASCAD variable-key database, which has a restricted number of training traces per key, recovering the key within 20 attack traces in comparison to 40 traces as required by the state-of-the-art CNN-based model with Plaintext feature extension (CNNP)-based model. During the profiling stage an attacker needs no additional knowledge of the implementation, such as the masking scheme or random mask values, only the ability to record the power consumption or electromagnetic field traces, plaintext/ciphertext and the key is needed. However, a two step training procedure is required. Additionally, no heuristic pre-processing is required in order to break the multiple masking countermeasures of the target implementation. Anh-Tuan Hoang, Neil Hanley, Ayesha Khalid, Dur-e-Shahwar Kundi, Máire O'Neill |
SECRYPT | 5 |
| 2022 | AxRLWE: A Multilevel Approximate Ring-LWE Co-Processor for Lightweight IoT ApplicationsabstractThis work presents a multilevel approximation exploration undertaken on the Ring-Learning-with-Errors (R-LWE)-based public-key cryptographic (PKC) schemes that belong to quantum-resilient cryptography algorithms. Among the various quantum-resilient cryptography schemes proposed in the currently running NIST’s post-quantum cryptography (PQC) standardization plan, the lattice-based learning-with-error (LWE) schemes have emerged as the most viable and preferred class for the Internet of Things (IoT) applications due to their compact area and memory footprint compared to other alternatives. However, compared to the classical schemes used today, R-LWE is much harder a challenge to fit on embedded IoT (end-node) devices, due to their stricter resource constraints (lower area, memory, and energy budgets) as well as their limited computational capabilities. To the best of our knowledge, this is the first endeavor exploring the inherent approximate nature of the LWE problem to undertake a multilevel approximate R-LWE (AxRLWE) architecture with respective security estimates opt for lightweight IoT devices. Undertaking AxRLWE on field-programmable gate arrays (FPGAs), we benchmarked a 64% area reduction cost compared to earlier accurate R-LWE designs at the cost of reduced quantum security. For the application-specific integrated circuits (ASICs) with 45-nm CMOS technology, AxRLWE was benchmarked to fit well within the same area budget of a lightweight ECC processor and consume a third of energy compared to special class of R-Binary LWE (R-BLWE) designs being proposed for an IoT, with a better security level. Dur-e-Shahwar Kundi, Ayesha Khalid, Song Bian 0001, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Internet Things J. | 5 |
| 2022 | Editorial Special Issue on Circuits and Systems for Emerging Computing ParadigmsabstractAS Dennard’s law is coming to an end, on-chip power consumption reduction and throughput improvement due to technology scaling pose serious challenges; workloads of today’s applications (such as AI, big data, and the IoT) have also reached extremely high levels of complex computation. Power dissipation has become the fundamental barrier to scale computing performance across all technology platforms. Computation at nanoscales requires innovative approaches. Shanshan Liu 0001, Bi Wu 0002, Ke Chen 0018, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | A High-Performance SIKE Hardware AcceleratorabstractSupersingular isogeny key encapsulation (SIKE) is a promising candidate in the NIST postquantum cryptography (PQC) standardization process, which has the smallest key lengths. It is the only isogeny-based cryptographic scheme in the NIST list that leverages the traditional elliptic curve cryptography (ECC) arithmetic; however, the high computational complexity is one of its limiting factors. In this work, we proposed a high-performance hardware architecture for the SIKE protocol. The architecture includes an improved multiplier based on the high-performance finite field multiplication (HFFM) algorithm which is 15%–20.7% faster than the previous multiplier based on the HFFM algorithm and a unified adder/subtractor with radix$3^{b}$. In addition, it also comprises an efficient scheduler strategy that decomposes all the functions of SIKE into finite field$F_{p}$and then effectively schedules through optimized multiplication chains for maximal performance. The proposed architecture is synthesized and implemented on Xilinx Virtex-7 FPGA for all the four variants of SIKE having security levels from 1 to 5 and achieved 2.6%–7.8% faster speeds as well as consumed less equivalent number of slices (ENS) than the state-of-the-art designs. In the comparison of area and time (AT), the proposed architecture is 14.2%–34.5% lower than the previous architecture. Ziying Ni, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | A Generic Dynamic Responding Mechanism and Secure Authentication Protocol for Strong PUFsabstractAs a lightweight hardware security primitive, physical unclonable functions (PUFs) can provide reliable identity authentication for devices of Internet of Things (IoTs) with limited resources. However, the delay-based PUF structures in authentication protocols have static responding behaviors, which make them vulnerable to modeling attacks. To address this issue, many complex PUF designs have been designed to increase the nonlinearity of their models. However, most of them can still be broken by modeling-based machine learning (ML) attacks. In this article, a dynamic responding mechanism for PUF designs to generate dynamic responses is proposed. Different from the concept of logically reconfigurable PUFs, the proposed mechanism does not rely on external inputs to provide reconfiguration signals. And different from the conventional PUF authentication protocols that use large-size linear feedback shift register (LFSR) to extend the master challenge, the proposed scheme uses internally generated dynamic signals to obfuscate the master challenge to generate multiple subchallenges. These subchallenges are then input to the underlying strong PUF to generate multibit dynamic responses. It can prevent an attacker from obtaining valid challenge-response pairs (CRPs) for the underlying PUF. A security authentication protocol is also proposed, the special authentication bit-string design can resist both conventional ML attacks and the latest covariance matrix adaptation evolution strategies (CMA-ES) variant. Yale Wang, Chenghua Wang, Chongyan Gu, Yijun Cui, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2021 | High-Performance Systolic Array Montgomery Multiplier for SIKEabstractIn theory, the speed of quantum computers is much faster than classical computers, which poses a threat to the Public Key Cryptography (PKC) that are currently in use. Post Quantum Cryptography (PQC) is a class of cryptography based on complex mathematical problems that are difficult to be attacked by quantum computers. The Supersingular Isogeny Key Encapsulation (SIKE) protocol is one of candidate algorithms for the US National Institute of Standards and Technology (NIST) PQC standardization process and survived to the Round 3. In this paper, we reconstruct the systolic array based Montgomery multiplier architecture for SIKE, using a three-stage pipeline that results in frequency improvement of 21.4%. The proposed multiplier consumed fewer DSP resources than the state-of-the-art SIKE designs and has a speed increase up to 12.7%. Ziying Ni, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2021 | Towards CRYSTALS-Kyber: A M-LWE Cryptoprocessor with Area-Time Trade-OffabstractCRYSTALS-Kyber is a quantum-resistant and promising lattice-based cryptography (LBC) in the finalists of the third round post-quantum cryptography (PQC) standardization, which is based on the hardness of Module-Learning with Errors (M-LWE). The variadic parameters make M-LWE obtain a more flexible security-performance trade-off than Ring-LWE. In this paper, we propose a M-LWE cryptoprocessor targeting CRYSTALS-Kyber with area-time trade-off for the first time. This balanced design includes a fast and low-cost Binomial Sampler and vector-polynomials multiplication structure based on pipelined decimation-in-frequency (DIF) based Number Theoretic Transform (NTT) technique. The M-LWE cryptoprocessor achieve 27,708 encryption operations per second using only 690 slices and 106,716 decryption operations per second using only 571 slices. Our proposed design achieved the lowest area-time product (ATP) with at least 2 χ performance improvement than the state-of-the-art LBC designs with a similar security level and complexity of polynomials. Kan Yao, Dur-e-Shahwar Kundi, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2021 | A Dynamic Highly Reliable SRAM-Based PUF Retaining Memory FunctionabstractIn this paper, a highly reliable SRAM based Physical Unclonable Function (PUF), which retains the memory function is proposed. The mismatch of NMOS is extracted during discharge process and amplified by the cross-coupled inverter to generate a response. At the beginning of the discharge process, the NMOSs are biased at sub-threshold region, which can improve the reliability and stability. The proposed PUF is designed in a 40nm CMOS process and each bit cell only consumes 4.98 μm2(3112F2). Post simulation shows that the bit error rate (BER) deterioration is 0.96% per 0.1V, 0.36% per 10° C with temperature variations from -40° C to 80° C and supply voltage variations from 0.9V to 1.3V. It achieves 1.8% native instability through the simulation. Meanwhile, the proposed PUF can retain memory function after a response is generated. Chenghua Wang, Chenggang Yan 0002, Yijun Cui, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 6 |
| 2021 | DTA-PUF: Dynamic Timing-aware Physical Unclonable Function for Resource-constrained DevicesabstractIn recent years, physical unclonable functions (PUFs) have gained a lot of attention as mechanisms for hardware-rooted device authentication. While the majority of the previously proposed PUFs derive entropy using dedicated circuitry, software PUFs achieve this from existing circuitry in a system. Such software-derived designs are highly desirable for low-power embedded systems as they require no hardware overhead. However, these software PUFs induce considerable processing overheads that hinder their adoption in resource-constrained devices. In this article, we propose DTA-PUF, a novel, software PUF design that exploits the instruction- and data-dependent dynamic timing behaviour of pipelined cores to provide a reliable challenge-response mechanism without requiring any extra hardware. DTA-PUF accepts sequences of instructions as an input challenge and produces an output response based on the manifested timing errors under specific over-clocked settings. To lower the required processing effort, we systematically select instruction sequences that maximise error-rate. The application to a post-layout pipelined floating-point unit, which is implemented in 45 nm process technology, demonstrates the effectiveness and practicability of our PUF design. Finally, DTA-PUF requires up to 50× fewer instructions than existing software processor PUF designs, limiting processing costs and resulting in up to 26% power savings. Ioannis Tsiokanos, Jack Miskelly, Chongyan Gu, Máire O'Neill, Georgios Karakonstantis |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2021 | A Modeling Attack Resistant Deception Technique for Securing Lightweight-PUF-Based AuthenticationabstractSilicon physical unclonable function (PUF) has emerged as a promising spoof-proof solution for low-cost device authentication. Due to practical constraints in preventing phishing through a public network or insecure communication channels, simple PUF-based authentication protocol with unrestricted queries and transparent responses is vulnerable to modeling and replay attacks. In this article, we present a modeling attack resistant PUF-based mutual authentication scheme to mitigate the practical limitations in applications where a resource-rich server authenticates a device with no strong restriction imposed on the type of PUF design or any additional protection on the binary channel used for the authentication. Our scheme uses an active deception protocol to prevent machine learning (ML) attacks on a device with a monolithic integration of a genuine strong PUF (SPUF), a fake PUF, a pseudorandom number generator (PRNG), a register, a binary counter, a comparator, and a simple controller. The hardware encapsulation makes the collection of challenge-response pairs (CRPs) easy for model building during enrollment but prohibitively time consuming upon device deployment through the same interface. A genuine server can perform a mutual authentication with the device using a combined fresh challenge contributed by both the server and the device. The message exchanged in clear cannot be manipulated by the adversary to derive unused authentic CRPs. The adversary will have to either wait for an impractically long time to collect enough real CRPs by directly querying the device or the ML model derived from the collected CRPs will be poisoned. The false PUF multiplexing is fortified against the prediction of waiting time by doubling the time penalty for every unsuccessful guess. Our implementation results on field-programmable gate array (FPGA) device and security analysis have corroborated the low hardware overheads and attack resistance of the proposed deception protocol. Chongyan Gu, Chip-Hong Chang, Weiqiang Liu 0001, Shichao Yu, Yale Wang, Máire O'Neill |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | A Secure Algorithm for Rounded Gaussian Sampling
Séamus Brannigan, Máire O'Neill, Ayesha Khalid, Ciara Rafferty |
CANS | 2 |
| 2020 | Security Analysis of Hardware Trojans on Approximate CircuitsabstractApproximate computing, for error-tolerant applications, provides trade-offs for computations to achieve improved speed and power performance. Approximate circuits, in particular approximate arithmetic circuits, directly affect the performance of a computing system. Hence, approximate circuit designs have been extensively studied. However, security issues of approximate circuits have been ignored. Moreover, hardware Trojans have been found in fabricated chips in manufacturing industry chains by untrusted foundries. Hardware Trojans could affect the functionality of approximate circuits under very rare circumstances with inconsiderable footprints. In this paper, hardware Trojan insertion methods based on signal transition probability are utilized to investigate and evaluate the security threats in approximate circuits. A approximate low-partor-adder (LOA) adder is utilized as an example and analyzed in the paper. The evaluation results show that with the increase of the number of approximation modules, the approximate LOA adder is more possible to be inserted hardware Trojans than the exact LOA adder. Yuqin Dou, Shichao Yu, Chongyan Gu, Máire O'Neill, Chenghua Wang, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | Programmable Ring Oscillator PUF Based on Switch MatrixabstractConfigurable ring oscillator (CRO) physical unclonable functions (PUFs) which can improve the uniqueness and reliability of conventional RO PUFs have been widely studied. Especially, the multiplier, XOR gate and tristate inverter based CRO PUFs can improve the uniqueness and reliability. However the efficiency is remain at the same level when compared with the conventional RO PUFs. In this paper, a programmable RO PUF (PRO PUF), which can be programmed to change the structure of a typical RO PUF, is proposed. The proposed PRO PUF design is implemented based on the switch matrix of an FPGA and can be programmed as a chained RO PUF or a random looped RO PUF. The proposed PRO PUF is implemented on Xilinx Spartan 6 FPGAs. Experimental results demonstrate that the proposed PRO PUF design has good uniqueness and reliability metrics as well as a high hardware efficiency. Yijun Cui, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2020 | AxMM: Area and Power Efficient Approximate Modular Multiplier for R-LWE CryptosystemabstractAmongst various Post-Quantum Cryptographic (PQC) schemes, Lattice-Based Cryptography (LBC) stands out as the most viable substitute to the classical cryptographic schemes due to its efficiency, versatility and solid foundations on hard mathematical problems. Ring Learning With Errors (R-LWE) is a Public Key Encryption (PKE) scheme of LBC, in which the modular polynomial multiplication in a ring is the main bottleneck in the realization of a practical resource-constraint design for the embedded IoT devices. This work explores novel Approximate Computing (AC) technique for the design of area/power efficient modular multiplier (so called AxMM) for R-LWE, exploiting the inherent approximate structure of the scheme. The proposed AxMM on 45nm ASIC library achieved an area and power reduction of 36% and 23%, respectively, along with a speed increase of 1.34× as compared to state-of-art smallest exact R-LWE modular multiplier. Dur-e-Shahwar Kundi, Song Bian 0001, Ayesha Khalid, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2020 | A Novel Feature Extraction Strategy for Hardware Trojan DetectionabstractHardware Trojans (HTs) are acknowledged as a significant emerging security concern in the IC industry resulting from the globalization of the semiconductor supply chain. Recently, taking advantage of the exponential growth in computing power, machine learning (ML) approaches such as neural networks (NNs) are being considered for HT detection. However, the circuit structure and components of an IC design are different from the data types in the ML models. To efficiently extract HT features from complex IC designs and utilize common ML-based detection approaches is challenging. In this paper, a novel HT feature extraction strategy based on gate-level circuit netlists is proposed to tackle the challenges. The HT features are extracted from the circuit topology rather than statistical analysis in previous research. A commonly utilized support vector machine (SVM)-based HT detection model is employed for data training and testing using the extracted features on HT benchmarks from both open-sourced library and HT generation platform to prove the feasibility and efficiency of the proposed HT feature extraction strategy. The detection results show high recall in nearly all tested benchmarks, achieving at most 97.7% recall on sequential Trojans and 84.8% on combinational ones. Shichao Yu, Chongyan Gu, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 4 |
| 2020 | Security in Approximate Computing and Approximate Computing for Security: Challenges and OpportunitiesabstractApproximate computing is an advanced computational technique that trades the accuracy of computation results for better utilization of system resources. It has emerged as a new preferable paradigm over traditional computing architectures for many applications where inaccurate results are acceptable. However, approximate computing also introduces security vulnerabilities mainly due to the fact that the uncertain and unpredictable intrinsic errors during approximate execution may be indistinguishable from malicious modification of the input data, the execution process, and the results. On the other hand, interestingly, approximate computing presents new opportunities to secure the system and the computation. Existing work on the security of approximate computing covers threat models, countermeasures, and evaluations but lacks a framework for analysis and comparison. In this article, we provide a classification of the state-of-the-art works in this research field, including threat models in approximate computing and promising security approaches using approximate computing. Open questions and potential future research directions are also discussed. Weiqiang Liu 0001, Chongyan Gu, Máire O'Neill, Gang Qu 0001, Paolo Montuschi, Fabrizio Lombardi |
Proc. IEEE | 3 |
| 2020 | High Performance Modular Multiplication for SIDHabstractThe latest research indicates that quantum computers will be realized in the near future. In theory, the computation speed of a quantum computer is much faster than current computers, which will pose a serious threat to current cryptosystems. Post-quantum cryptography (PQC) is a class of cryptography based on underlying mathematical problems that are considered infeasible to crack even with access to a quantum computer. The supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol is a new post-quantum cryptosystem, which offers advantages in reduced secret key length and attack resistance. SIDH is the basis of the supersingular isogeny key encapsulation (SIKE) protocol, which is in the second round of the U.S. National Institute of Standards and Technology (NIST) PQC standardization process. In this article, we propose a new modular multiplication algorithm and a new interleaved hardware architecture for SIDH. Performance results for the proposed modular multiplier using four parameter sets for the prime, p that correspond to the SIKE Round 2 parameter sets show significant advantages in speed. Weiqiang Liu 0001, Ziying Ni, Jian Ni, Ciara Rafferty, Máire O'Neill |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Fast DRAM PUFs on Commodity DevicesabstractIntrinsic physical unclonable functions (PUFs), which derive hardware identifiers from components already present in a system without modification, are an appealing way to add a layer of hardware rooted security into a system. This is evidenced by the fact that the majority of PUF designs in commercial use today are intrinsic. However, as each intrinsic PUF design is reliant on specific hardware their use is limited to a subset of systems. It is therefore desirable to have practical intrinsic PUF designs for as wide a range of underlying hardware as possible. Most intrinsic PUF designs to date have used memory as the entropy source, with the most well studied type being based on SRAM. More recently designs based on DRAM have been proposed, an appealing prospect considering the ubiquity of that technology. While previous research has demonstrated that entropy can be extracted from DRAM there has not yet been a substantive demonstration of such a PUF operating in real-time on a commodity system. In this article, we present a novel set of algorithms for deriving PUF responses in-runtime from DRAM by altering timing parameters using only software. These algorithms reduce the critical period of system disruption by 96% from 88 ms to 3 ms on average compared to existing designs. We present a large scale dataset derived from 1824 DRAM chips characterized using the proposed design on commodity off-the-shelf desktop hardware running a Linux OS. An analysis of the data shows that in addition to the speed improvements the proposed design shows near ideal (>44%) uniqueness and good (>88%) reliability. Jack Miskelly, Máire O'Neill |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Theoretical Analysis of Delay-Based PUFs and Design Strategies for ImprovementabstractDelay-based physical unclonable function (PUF) designs use the random delay differences in circuit transmission to extract response. In the existing PUF designs, there are few studies on investigating the link between process variation and PUF performance. The experimental data can reflect the performance of the new design to a certain extent, but lack of theoretical analysis to provide thorough information. In this paper, a theoretical model for delay-based PUF designs is proposed. An analysis of the delay-based PUF improvements by existing design strategies is also investigated. Moreover, a guidance to develop and improve future delay-based PUF designs using the proposed theoretical model is also given in this paper. Yale Wang, Chenghua Wang, Chongyan Gu, Yijun Cui, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2019 | Multi-Incentive Delay-Based (MID) PUFabstractThis paper proposes a new PUF, namely Multi-incentive Delay-based PUF (MID PUF), which utilizes the fast carry logic (FCL) of Field Programmable Gate Arrays (FPGAs). The proposed MID PUF is completely and efficiently implemented in XOR gates of FCLs. Compared to other single signal excited PUF designs, e.g. Arbiter PUF, multiple excitations are applied on the same delay line to produce multiple outputs. To the authors' best knowledge, this is the first strong PUF based on only FCLs. The proposed MID PUF is implemented on Xilinx Spartan-6 XC6SLX9 FPGAs and a reliability experiment is carried out under the operating temperature in a range of 0°C~70° C. The experimental results show that the proposed MID PUF has a high uniqueness and reliability performance, as well as low hardware consumption. Due to its advantages in both hardware efficiency and PUF metrics, the proposed MID PUF is promising for low-cost security applications on FPGAs. Zhengran Zhang, Chongyan Gu, Yijun Cui, Chuan Zhang 0001, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2019 | A Theoretical Model to Link Uniqueness and Min-Entropy for PUF EvaluationsabstractPhysical unclonable functions (PUFs) are security primitives that enable the extraction of digital identifiers from electronic devices, based on the inherent silicon process variations between devices which occur during the manufacturing process. Due to the intrinsic and lightweight nature of a PUF, they have been proposed to provide security at a low cost for many applications, in particular for the internet of things (IoT). Many metrics have been proposed to evaluate the security and performance of PUF architectures, two of which are uniqueness and min-entropy. The uniqueness of a PUF response evaluates its ability to differentiate between different physical devices, while the min-entropy estimation is a measure of how much uncertainty the PUF response contains. The min-entropy is a lower-bound of real entropy. When the uniqueness of a PUF design is close to the optimal, it is unclear if this also implies that the design has a significantly high entropy; hence it would be useful to ascertain the minimum uniqueness required to achieve a given entropy. To date, a thorough investigation of the relationship between uniqueness and entropy for PUF designs has not been conducted. In this paper, this relationship between the uniqueness and entropy is explored, and for the first time, to the authors' knowledge, the relationship between them is modeled. To verify this model, both simulated and hardware-based experimental results are performed, with a test-bed containing 184 Xilinx Artix-7 FPGA based Basys3 boards providing a large data set for granular results. The experimental results demonstrate that the proposed model accurately estimates the relationship between uniqueness and min-entropy, with both the theoretical analysis and software simulations closely matching the experimental results. Chongyan Gu, Weiqiang Liu 0001, Neil Hanley, Robert Hesselbarth, Máire O'Neill |
IEEE Trans. Computers | 5 |
| 2019 | Optimized Modular Multiplication for Supersingular Isogeny Diffie-HellmanabstractRecent progress in quantum physics shows that quantum computers may be a reality in the not too distant future. Post-quantum cryptography (PQC) refers to cryptographic schemes that are based on hard problems which are believed to be resistant to attacks from quantum computers. The supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol shows promising security properties among various post-quantum cryptosystems that have been proposed. In this paper, we propose two efficient modular multiplication algorithms with special primes that can be used in SIDH key exchange protocol. Hardware architectures for the two proposed algorithms are also proposed. The hardware implementations are provided and compared with the original modular multiplication algorithm. The results show that the proposed finite field multiplier is over 6.79 times faster than the original multiplier in hardware. Moreover, the SIDH hardware/software codesign implementation using the proposed FFM2 hardware is over 31 percent faster than the best SIDH software implementation. Weiqiang Liu 0001, Jian Ni, Zhe Liu 0001, Máire O'Neill |
IEEE Trans. Computers | 5 |
| 2019 | XOR-Based Low-Cost Reconfigurable PUFs for IoT SecurityabstractWith the rapid development of the Internet of Things (IoT), security has attracted considerable interest. Conventional security solutions that have been proposed for the Internet based on classical cryptography cannot be applied to IoT nodes as they are typically resource-constrained. A physical unclonable function (PUF) is a hardware-based security primitive and can be used to generate a key online or uniquely identify an integrated circuit (IC) by extracting its internal random differences using so-called challenge-response pairs (CRPs). It is regarded as a promising low-cost solution for IoT security. A logic reconfigurable PUF (RPUF) is highly efficient in terms of hardware cost. This article first presents a new classification for RPUFs, namely circuit-based RPUF (C-RPUF) and algorithm-based RPUF (A-RPUF); two Exclusive OR (XOR)-based RPUF circuits (an XOR-based reconfigurable bistable ring PUF (XRBR PUF) and an XOR-based reconfigurable ring oscillator PUF (XRRO PUF)) are proposed. Both the XRBR and XRRO PUFs are implemented on Xilinx Spartan-6 field-programmable gate arrays (FPGAs). The implementation results are compared with previous PUF designs and show good uniqueness and reliability. Compared to conventional PUF designs, the most significant advantage of the proposed designs is that they are highly efficient in terms of hardware cost. Moreover, the XRRO PUF is the most efficient design when compared with previous RPUFs. Also, both the proposed XRRO and XRBR PUFs require only 12.5% of the hardware resources of previous bitstable ring PUFs and reconfigurable RO PUFs, respectively, to generate a 1-bit response. This confirms that the proposed XRBR and XRRO PUFs are very efficient designs with good uniqueness and reliability. Weiqiang Liu 0001, Lei Zhang 0089, Zhengran Zhang, Chongyan Gu, Chenghua Wang, Máire O'Neill, Fabrizio Lombardi |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2019 | Optimized Schoolbook Polynomial Multiplication for Compact Lattice-Based Cryptography on FPGAabstractLattice-based cryptography (LBC) is one of the most promising classes of post-quantum cryptography (PQC) that is being considered for standardization. This brief proposes an optimized schoolbook polynomial multiplication (SPM) for compact LBC. We exploit the symmetric nature of Gaussian noise for bit reduction. Additionally, a single field-programmable gate array (FPGA) DSP block is used for two parallel multiplication operations per clock cycle. These optimizations enable a significant 2.2× speedup along with reduced resources for dimension n = 256. The overall efficiency (throughput per slice) is 1.28× higher than the conventional SPM, as well as contributing to a more compact LBC system compared to previously reported designs. The results targeting the FPGA platform show that the proposed design can achieve high hardware efficiency with reduced hardware area costs. Weiqiang Liu 0001, Sailong Fan, Ayesha Khalid, Ciara Rafferty, Máire O'Neill |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | A machine learning attack resistant multi-PUF design on FPGAabstractCurrent approaches for building physical unclonable function (PUF) designs resistant to machine learning attacks often suffer from large resource overhead and are typically difficult to implement on field programmable gate arrays (FPGAs). In this paper we propose a new arbiter-based multi-PUF (MPUF) design that utilises a Weak PUF to obfuscate the challenges to a Strong PUF and is harder to model than the conventional arbiter PUF using machine learning attacks. The proposed PUF design shows a greater resistance to attacks, which have been successfully applied to other Arbiter PUFs. A mathematical model is presented to analyse the complexity and obfuscation properties of the proposed PUF design. Moreover, we show that it is feasible to implement the proposed MPUF design on a Xilinx Artix-7 FPGA, and that it achieves a good uniqueness result of 40.60 % and uniformity of 37.03 %, which significantly improves over previous work into multi-PUF designs. Qingqing Ma, Chongyan Gu, Neil Hanley, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill |
ASP-DAC | 6 |
| 2018 | Physical Protection of Lattice-Based Cryptography: Challenges and SolutionsabstractThe impending realization of scalable quantum computers will have a significant impact on today's security infrastructure. With the advent of powerful quantum computers public key cryptographic schemes will become vulnerable to Shor's quantum algorithm, undermining the security current communications systems. Post-quantum (or quantum-resistant) cryptography is an active research area, endeavoring to develop novel and quantum resistant public key cryptography. Amongst the various classes of quantum-resistant cryptography schemes, lattice-based cryptography is emerging as one of the most viable options. Its efficient implementation on software and on commodity hardware has already been shown to compete and even excel the performance of current classical security public-key schemes. This work discusses the next step in terms of their practical deployment, i.e., addressing the physical security of lattice-based cryptographic implementations. We survey the state-of-the-art in terms of side channel attacks (SCA), both invasive and passive attacks, and proposed countermeasures. Although the weaknesses exposed have led to countermeasures for these schemes, the cost, practicality and effectiveness of these on multiple implementation platforms, however, remains under-studied. Ayesha Khalid, Tobias Oder, Felipe Valencia, Máire O'Neill, Tim Güneysu, Francesco Regazzoni 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | Compact, Scalable, and Efficient Discrete Gaussian Samplers for Lattice-Based CryptographyabstractLattice-based cryptography, one of the leading candidates for post-quantum security, relies heavily on discrete Gaussian samplers to provide necessary uncertainty, obfuscating computations on secret information. For reconfigurable hardware, the cumulative distribution table (CDT) scheme has previously been shown to achieve the highest throughput and the smallest resource utilisation, easily outperforming other existing samplers. However, the CDT sampler does not scale well. In fact, for large parameters, the lookup tables required are far too large to be practically implemented. This research proposes a hierarchy of multiple smaller samplers, extending the Gaussian convolution lemma to compute optimal parameters, where the individual samplers require much smaller lookup tables. A large range of parameter sets, covering encryption, signatures, and key exchange are evaluated. Hardware-optimised parameters are formulated and a practical implementation on Xilinx Artix-7 FPGA device is realised. The proposed sampling designs demonstrate promising performance on reconfigurable hardware, even for large parameters, that were otherwise thought infeasible. Ayesha Khalid, James Howe, Ciara Rafferty, Francesco Regazzoni 0001, Máire O'Neill |
ISCAS | 5 |
| 2018 | Design and Optimization of Modular Multiplication for SIDHabstractRecent progress on quantum physics shows that quantum computers may be a reality in the not too distant future. Based on new mathematical hard problems, post-quantum cryptography (PQC) has been studied to make sure the attacks from quantum computers can be resistant. The latest supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol shows promising security properties among various post-quantum cryptosystems. In this paper, we propose an improved modular multiplication algorithm with special primes that can be used in SIDH key exchange protocol. Both software and hardware implementations are provided and compared with original modular multiplication algorithm. The results show that the software results of improved algorithm can be 24% faster than the original software implementation, while the hardware implementation based on the proposed hardware architecture can be 6 times faster than previous hardware implementation. Jian Ni, Weiqiang Liu 0001, Zhe Liu 0001, Máire O'Neill |
ISCAS | 5 |
| 2018 | Design of Majority Logic (ML) Based Approximate Full AddersabstractAs a new paradigm in the nanoscale technologies, approximate computing enables error tolerance in the computational process; it has also emerged as a low power design methodology for arithmetic circuits. Majority logic (ML) is applicable to many emerging technologies and its basic building block (the 3-input majority voter) has been extensively used in digital circuit design. In this paper, we propose the design of a one-bit approximate full adder based on majority logic. Furthermore, multi-bit approximate full adders are also proposed and studied; the application of these designs to quantum-dot cellular automata (QCA) is also presented as an example. The designs are evaluated using hardware metrics (including delay and area) as well as error metrics. Compared with other circuits found in the technical literature, the optimal designs are found to offer superior performance. Weiqiang Liu 0001, Emma McLarnon, Máire O'Neill, Fabrizio Lombardi |
ISCAS | 4 |
| 2018 | On Practical Discrete Gaussian Samplers for Lattice-Based CryptographyabstractLattice-based cryptography is one of the most promising branches of quantum resilient cryptography, offering versatility and efficiency. Discrete Gaussian samplers are a core building block in most, if not all, lattice-based cryptosystems, and optimised samplers are desirable both for high-speed and low-area applications. Due to the inherent structure of existing discrete Gaussian sampling methods, lattice-based cryptosystems are vulnerable to side-channel attacks, such as timing analysis. In this paper, the first comprehensive evaluation of discrete Gaussian samplers in hardware is presented, targeting FPGA devices. Novel optimised discrete Gaussian sampler hardware architectures are proposed for the main sampling techniques. An independent-time design of each of the samplers is presented, offering security against side-channel timing attacks, including the first proposed constant-time Bernoulli, Knuth-Yao, and discrete Ziggurat sampler hardware designs. For a balanced performance, the Cumulative Distribution Table (CDT) sampler is recommended, with the proposed hardware CDT design achieving a throughput of 59.4 million samples per second for encryption, utilising just 43 slices on a Virtex 6 FPGA and 16.3 million samples per second for signatures with 179 slices on a Spartan 6 device. James Howe, Ayesha Khalid, Ciara Rafferty, Francesco Regazzoni 0001, Máire O'Neill |
IEEE Trans. Computers | 5 |
| 2017 | FPGA-based strong PUF with increased uniqueness and entropy propertiesabstractPhysical unclonable functions (PUFs), are a type of physical security primitive which enable identification and authentication of hardware devices, such as field programmable gate arrays (FPGAs) and application specific integrated circuits (ASICs). Arbiter PUFs were the first proposed Strong PUF and are also widely studied. However, these designs often suffer from poor uniqueness and reliability characteristics leaving them vulnerable to modeling attacks, as well as being difficult to implement on FPGAs due to the physical layout restrictions. Some more recent designs based around non-linear voltage transfer characteristics, or non-linear currents improve the resistance against modeling attacks. However they can only be implemented on ASICs due to their voltage/current requirements. To address this problem, we propose a new PUF circuit that offers a significantly higher theoretical entropy than the traditional Arbiter PUF construction, and which is specifically designed for FPGAs. The proposed work is verified on a low-cost Nexys4 board which contains a Xilinx Artix-7 FPGA fabricated at 28nm. The experimental results give a uniqueness of 20 %, considerably higher than the reported 9 % of a traditional Arbiter PUF design, and an expected reliability of ≈ 96% over an environmental temperature range of 0° C to 75° C, with a reliability of ≈ 92 % with ±10 % variation in supply voltage. Chongyan Gu, Neil Hanley, Máire O'Neill |
ISCAS | 3 |
| 2017 | Compact and provably secure lattice-based signatures in hardwareabstractLattice-based cryptography is a quantum-safe alternative to existing classical asymmetric cryptography, such as RSA and ECC, which may be vulnerable to future attacks in the event of the creation of a viable quantum computer. The efficiency of lattice-based cryptography has improved over recent years, but there has been relatively little investigation into hardware designs of digital signature schemes. In this paper, the first hardware design of the provably secure Ring-LWE digital signature scheme, Ring-TESLA, is presented, targeting a Xilinx Spartan-6 FPGA. The results better compactness of all previous lattice-based digital signature schemes in hardware, and can achieve between 104-785 signatures and 102-776 verifications per second. James Howe, Ciara Rafferty, Ayesha Khalid, Máire O'Neill |
ISCAS | 4 |
| 2017 | XOR gate based low-cost configurable RO PUFabstractA Physical Unclonable Function (PUF) is often used to uniquely identify an integrated circuit by extracting its internal random differences using so-called Challenge Response Pairs (CRPs). As CRPs include unique information about the underlying hardware variations, PUF design is a promising approach to provide authentication and IP-protection capabilities. In this paper, an XOR-gate-based configurable Ring Oscillator (RO) PUF (denoted as XCRO PUF) is presented. This XCRO PUF can generate more CRPs compared with state-of-the-art PUF designs by using the same number of configurable logic blocks (CLBs) in an FPGA implementation. This design is implemented in the Xilinx Spartan-6 XC6SLX9 FPGAs with fixed locations for the XCROs (placed within a ring to improve its uniqueness). The XCRO PUF shows better uniqueness and reliability than other PUF designs. Moreover, a XCRO PUF consumes only 12.5% of the hardware resources to generate a 1-bit response compared with other CRO PUFs implemented in FPGA. Lei Zhang 0089, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
ISCAS | 4 |
| 2017 | GLITCH: A Discrete Gaussian Testing Suite for Lattice-based CryptographyabstractLattice-based cryptography is one of the most promising areas within post-quantum cryptography, and offers versatile, efficient, and high performance security services. The aim of this paper is to verify the correctness of the discrete Gaussian sampling component, one of the most important modules within lattice-based cryptography. In this paper, the GLITCH software test suite is proposed, which performs statistical tests on discrete Gaussian sampler outputs. An incorrectly operating sampler, for example due to hardware or software errors, has the potential to leak secret-key information and could thus be a potential attack vector for an adversary. Moreover, statistical test suites are already common for use in pseudo-random number generators (PRNGs), and as lattice-based cryptography becomes more prevalent, it is important to develop a method to test the correctness and randomness for discrete Gaussian sampler designs. Additionally, due to the theoretical requirements for the discrete Gaussian distribution within lattice-based cryptography, certain statistical tests for distribution correctness become unsuitable, therefore a number of tests are surveyed. The final GLITCH test suite provides 11 adaptable statistical analysis tests that assess the exactness of a discrete Gaussian sampler, and which can be used to verify any software or hardware sampler design. James Howe, Máire O'Neill |
SECRYPT | 2 |
| 2017 | Evaluation of Large Integer Multiplication Methods on HardwareabstractMultipliers requiring large bit lengths have a major impact on the performance of many applications, such as cryptography, digital signal processing (DSP) and image processing. Novel, optimised designs of large integer multiplication are needed as previous approaches, such as schoolbook multiplication, may not be as feasible due to the large parameter sizes. Parameter bit lengths of up to millions of bits are required for use in cryptography, such as in lattice-based and fully homomorphic encryption (FHE) schemes. This paper presents a comparison of hardware architectures for large integer multiplication. Several multiplication methods and combinations thereof are analysed for suitability in hardware designs, targeting the FPGA platform. In particular, the first hardware architecture combining Karatsuba and Comba multiplication is proposed. Moreover, a hardware complexity analysis is conducted to give results independent of any particular FPGA platform. It is shown that hardware designs of combination multipliers, at a cost of additional hardware resource usage, can offer lower latency compared to individual multiplier designs. Indeed, the proposed novel combination hardware design of the Karatsuba-Comba multiplier offers lowest latency for integers greater than 512 bits. For large multiplicands, greater than 16,384 bits, the hardware complexity analysis indicates that the NTT-Karatsuba-Schoolbook combination is most suitable. Ciara Rafferty, Máire O'Neill, Neil Hanley |
IEEE Trans. Computers | 2 |
| 2017 | Improved Reliability of FPGA-Based PUF Identification Generator DesignabstractPhysical unclonable functions (PUFs), a form of physical security primitive, enable digital identifiers to be extracted from devices, such as field programmable gate arrays (FPGAs). Many PUF implementations have been proposed to generate these unique n -bit binary strings. However, they often offer insufficient uniqueness and reliability when implemented on FPGAs and can consume excessive resources. To address these problems, in this article we present an efficient, lightweight, and scalable PUF identification (ID) generator circuit that offers a compact design with good uniqueness and reliability properties and is specifically designed for FPGAs. A novel post-characterisation methodology is also proposed that improves the reliability of a PUF without the need for any additional hardware resources. Moreover, the proposed post-characterisation method can be generally used for any FPGA-based PUF designs. The PUF ID generator consumes 8.95% of the hardware resources of a low-cost Xilinx Spartan-6 LX9 FPGA and 0.81% of a Xilinx Artix-7 FPGA. Experimental results show good uniqueness, reliability, and uniformity with no occurrence of bit-aliasing. In particular, the reliability of the PUF is close to 100% over an environmental temperature range of 25°C to 70°C with ± 10% variation in the supply voltage. Chongyan Gu, Neil Hanley, Máire O'Neill |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2016 | Standard lattices in hardwareabstractLattice-based cryptography has gained credence recently as a replacement for current public-key cryptosystems, due to its quantum-resilience, versatility, and relatively low key sizes. To date, encryption based on the learning with errors (LWE) problem has only been investigated from an ideal lattice standpoint, due to its computation and size efficiencies. However, a thorough investigation of standard lattices in practice has yet to be considered. Standard lattices may be preferred to ideal lattices due to their stronger security assumptions and less restrictive parameter selection process. James Howe, Ciara Rafferty, Máire O'Neill, Francesco Regazzoni 0001, Tim Güneysu, K. Beeden |
DAC | 3 |
| 2016 | Time-independent discrete Gaussian sampling for post-quantum cryptographyabstractAs the development of a viable quantum computer nears, existing widely used public-key cryptosystems, such as RSA, will no longer be secure. Thus, significant effort is being invested into post-quantum cryptography (PQC). Lattice-based cryptography (LBC) is one such promising area of PQC, which offers versatile, efficient, and high performance security services. However, the vulnerabilities of these implementations against side-channel attacks (SCA) remain significantly understudied. Most, if not all, lattice-based cryptosystems require noise samples generated from a discrete Gaussian distribution, and a successful timing analysis attack can render the whole cryptosystem broken, making the discrete Gaussian sampler the most vulnerable module to SCA. This research proposes countermeasures against timing information leakage with FPGA-based designs of the CDT-based discrete Gaussian samplers with constant response time, targeting encryption and signature scheme parameters. The proposed designs are compared against the state-of-the-art and are shown to significantly outperform existing implementations. For encryption, the proposed sampler is 9× faster in comparison to the only other existing time-independent CDT sampler design. For signatures, the first time-independent CDT sampler in hardware is proposed. Ayesha Khalid, James Howe, Ciara Rafferty, Máire O'Neill |
FPT | 4 |
| 2016 | Live demonstration: An automatic evaluation platform for physical unclonable function testabstractPUF is a security primitive that exploits the fact that no two ICs are exactly the same. To verify a new PUF design, several metrics including uniqueness, reliability, and randomness must be evaluated, which requires various resources and a long set-up time. In this live demonstration, we have developed an automatically evaluation platform for the PUF design. To the authors' best knowledge, this is the first automatic evaluation platform for the PUF test. The evaluation platform can be used for both FPGA and ASCI PUF testing. Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 4 |
| 2016 | Low-cost configurable ring oscillator PUF with improved uniquenessabstractThe physical unclonable function (PUF) produces die-unique responses and is regarded as an emerging security primitive that can be used for authentication of devices. The complexity of a conventional PUF design based on a ring oscillator (RO) is rather high, so limiting its use in many applications. The configurable ring oscillator (CRO) PUF has been advocated as a possible solution to this issue. In this paper, a low hardware complexity CRO PUF design with an enhanced capability to generate a large number of bit responses is proposed; only an inverter and a multiplexer are used in each delay unit. The responses are generated by considering the variation due to fabrication of the logic gates and wires in the CROs. A novel comparison strategy is proposed for the generation of the responses. The proposed PUF design is implemented on Xilinx Spartan-6 FPGAs. These results show that the proposed CRO PUF design has good uniqueness; moreover, it is also robust in its operation for the temperature range of -25°C~85°C. Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
ISCAS | 5 |
| 2016 | Optimised Multiplication Architectures for Accelerating Fully Homomorphic EncryptionabstractLarge integer multiplication is a major performance bottleneck in fully homomorphic encryption (FHE) schemes over the integers. In this paper two optimised multiplier architectures for large integer multiplication are proposed. The first of these is a low-latency hardware architecture of an integer-FFT multiplier. Secondly, the use of low Hamming weight (LHW) parameters is applied to create a novel hardware architecture for large integer multiplication in integer-based FHE schemes. The proposed architectures are implemented, verified and compared on the Xilinx Virtex-7 FPGA platform. Finally, the proposed implementations are employed to evaluate the large multiplication in the encryption step of FHE over the integers. The analysis shows a speed improvement factor of up to 26.2 for the low-latency design compared to the corresponding original integer-based FHE software implementation. When the proposed LHW architecture is combined with the low-latency integer-FFT accelerator to evaluate a single FHE encryption operation, the performance results show that a speed improvement by a factor of approximately 130 is possible. Xiaolin Cao, Ciara Rafferty, Máire O'Neill, Elizabeth O'Sullivan, Neil Hanley |
IEEE Trans. Computers | 3 |
| 2016 | Design and Analysis of Inexact Floating-Point AddersabstractPower has become a key constraint in nanoscale integrated circuit design due to the increasing demands for mobile computing and higher integration density. As an emerging computational paradigm, an inexact circuit offers a promising approach to significantly reduce both dynamic and static power dissipation for error-tolerant applications. In this paper, an inexact floating-point adder is proposed by approximately designing an exponent subtractor and mantissa adder. Related operations such as normalization and rounding are also dealt with in terms of inexact computing. An upper bound error analysis for the average case is presented to guide the inexact design; it shows that the inexact floating-point adder design is dependent on the application data range. High dynamic range images are then processed using the proposed inexact floating-point adders to show the validity of the inexact design; comparison results show that the proposed inexact floating-point adders can improve the power consumption and power-delay product by 29.98 and 39.60 percent, respectively. Weiqiang Liu 0001, Linbin Chen, Chenghua Wang, Máire O'Neill, Fabrizio Lombardi |
IEEE Trans. Computers | 4 |
| 2015 | On the Security of Balanced Encoding Countermeasures
Yoo-Seung Won, Philip Hodgers, Máire O'Neill, Dong-Guk Han |
CARDIS | 3 |
| 2015 | Ultra-compact and robust FPGA-based PUF identification generatorabstractPhysically Unclonable Functions (PUFs), exploit inherent manufacturing variations and present a promising solution for hardware security. They can be used for key storage, authentication and ID generations. Low power cryptographic design is also very important for security applications. However, research to date on digital PUF designs, such as Arbiter PUFs and RO PUFs, is not very efficient. These PUF designs are difficult to implement on Field Programmable Gate Arrays (FPGAs) or consume many FPGA hardware resources. In previous work, a new and efficient PUF identification generator was presented for FPGA. The PUF identification generator is designed to fit in a single slice per response bit by using a 1-bit PUF identification generator cell formed as a hard-macro. In this work, we propose an ultra-compact PUF identification generator design. It is implemented on ten low-cost Xilinx Spartan-6 FPGA LX9 microboards. The resource utilization is only 2.23%, which, to the best of the authors' knowledge, is the most compact and robust FPGA-based PUF identification generator design reported to date. This PUF identification generator delivers a stable range of uniqueness of around 50% and good reliability between 85% and 100%. Chongyan Gu, Máire O'Neill |
ISCAS | 2 |
| 2015 | Pre-processing power traces to defeat random clocking countermeasuresabstractWe describe a pre-processing correlation attack on an FPGA implementation of AES, protected with a random clocking countermeasure that exhibits complex variations in both the location and amplitude of the power consumption patterns of the AES rounds. It is demonstrated that the merged round patterns can be pre-processed to identify and extract the individual round amplitudes, enabling a successful power analysis attack. We show that the requirement of the random clocking countermeasure to provide a varying execution time between processing rounds can be exploited to select a sub-set of data where sufficient current decay has occurred, further improving the attack. In comparison with the countermeasure's estimated security of 3 million traces from an integration attack, we show that through application of our proposed techniques that the countermeasure can now be broken with as few as 13k traces. Philip Hodgers, Neil Hanley, Máire O'Neill |
ISCAS | 3 |
| 2015 | RO PUF design in FPGAs with new comparison strategiesabstractA Physical Unclonable Function (PUF) can be used to provide authentication of devices by producing die-unique responses. In PUFs based on ring oscillators (ROs), the responses are derived from the oscillation frequencies of the ROs. However, RO PUFs can be vulnerable to attack due to the frequency distribution characteristics of the RO arrays. In this paper, in order to improve the design of RO PUFs for FPGA devices, the frequencies of RO arrays implemented on a large number of FPGA chips are statistically analyzed. Three RO frequency distribution (ROFD) characteristics are observed and discussed. Based on these ROFD characteristics, two RO comparison strategies are proposed that can be used to improve the design of RO PUFs. It is found that the symmetrical RO comparison strategy has the highest entropy density. Weiqiang Liu 0001, Chenghua Wang, Yijun Cui, Máire O'Neill |
ISCAS | 5 |
| 2015 | Privacy region protection for H.264/AVC with enhanced scrambling effect and a low bitrate overhead
Máire O'Neill, Fatih Kurugollu, Elizabeth O'Sullivan |
Signal Process. Image Commun. | 2 |
| 2015 | Practical Lattice-Based Digital Signature SchemesabstractDigital signatures are an important primitive for building secure systems and are used in most real-world security protocols. However, almost all popular signature schemes are either based on the factoring assumption (RSA) or the hardness of the discrete logarithm problem (DSA/ECDSA). In the case of classical cryptanalytic advances or progress on the development of quantum computers, the hardness of these closely related problems might be seriously weakened. A potential alternative approach is the construction of signature schemes based on the hardness of certain lattice problems that are assumed to be intractable by quantum computers. Due to significant research advancements in recent years, lattice-based schemes have now become practical and appear to be a very viable alternative to number-theoretic cryptography. In this article, we focus on recent developments and the current state of the art in lattice-based digital signatures and provide a comprehensive survey discussing signature schemes with respect to practicality. Additionally, we discuss future research areas that are essential for the continued development of lattice-based cryptography. James Howe, Thomas Pöppelmann, Máire O'Neill, Elizabeth O'Sullivan, Tim Güneysu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Introduction for Embedded Platforms for Cryptography in the Coming Decadeabstracteditorial Free Access Share on Introduction for Embedded Platforms for Cryptography in the Coming Decade Editors: Patrick Schaumont Virginia Tech, USA Virginia Tech, USAView Profile , Maire O'Neill Queen's University Belfast, United Kingdom Queen's University Belfast, United KingdomView Profile , Tim Güneysu Ruhr University Bochum, Germany Ruhr University Bochum, GermanyView Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 14Issue 3May 2015 Article No.: 40pp 1–3https://doi.org/10.1145/2745710Published:21 April 2015Publication History 2citation284DownloadsMetricsTotal Citations2Total Downloads284Last 12 Months15Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Patrick Schaumont, Máire O'Neill, Tim Güneysu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2014 | A unique and robust single slice FPGA identification generatorabstractIn this paper, a new field-programmable gate array (FPGA) identification generator circuit is introduced based on physically unclonable function (PUF) technology. The new identification generator is able to convert flip-flop delay path variations to unique n-bit digital identifiers (IDs), while requiring only a single slice per ID bit by using 1-bit ID cells formed as hard-macros. An exemplary 128-bit identification generator is implemented on ten Xilinx Spartan-6 FPGA devices. Experimental results show an uniqueness of 48.52%, and reliability of 92.41% over a 25°C to 70°C temperature range and 10% fluctuation in supply voltage. Chongyan Gu, Julian P. Murphy, Máire O'Neill |
ISCAS | 3 |
| 2014 | Practical homomorphic encryption: A surveyabstractCloud computing technology has rapidly evolved over the last decade, offering an alternative way to store and work with large amounts of data. However data security remains an important issue particularly when using a public cloud service provider. The recent area of homomorphic cryptography allows computation on encrypted data, which would allow users to ensure data privacy on the cloud and increase the potential market for cloud computing. A significant amount of research on homomorphic cryptography appeared in the literature over the last few years; yet the performance of existing implementations of encryption schemes remains unsuitable for real time applications. One way this limitation is being addressed is through the use of graphics processing units (GPUs) and field programmable gate arrays (FPGAs) for implementations of homomorphic encryption schemes. This review presents the current state of the art in this promising new area of research and highlights the interesting remaining open problems. Ciara Rafferty, Máire O'Neill, Elizabeth O'Sullivan, Yarkin Doröz, Berk Sunar |
ISCAS | 2 |
| 2013 | Privacy region protection for H.264/AVC by encrypting the intra prediction modes without drift error in I framesabstractWhile video surveillance systems have become ubiquitous in our daily life, this has also brought some serious privacy concerns. Therefore, recent research in this area has included a focus on privacy region protection. Existing video scrambling techniques are being applied to specific regions of interest in a video while the background is left unchanged. In this paper, a new method involving the encryption of intra prediction modes (IPM) is proposed for privacy region protection without drift error in I frames. Compared with a previous technique that uses encryption of IPM, the proposed method offers savings in the bitrate overhead. To enhance the scrambling effect for the privacy region, the proposed method can support the combination of IPM with encryption of the sign bits of nonzero quantized transform coefficients in the privacy region. Experimental results and analysis based on H.264/AVC were carried out and verify the effectiveness of the proposed method. Máire O'Neill, Fatih Kurugollu |
ICASSP | 2 |
| 2013 | Power analysis attack of QCA circuits: A case study of the Serpent cipherabstractQuantum-dot cellular automata (QCA) technology is an attractive alternative to CMOS for future digital designs. A powerful attack based on power analysis has become a significant threat to the security of CMOS cryptographic circuits. As there is no current flow in QCA, the power consumption of a QCA circuit is extremely low compared to its CMOS counterpart. Therefore, in this paper an investigation is carried out to ascertain if QCA circuits could be immune to power analysis attacks based on a case study of the Serpent cipher. In comparison to a previous design, the proposed QCA implementation of a sub-module of the Serpent cipher is more efficient in terms of complexity, area and latency. By using an upper bound power model, the first power analysis attack of a QCA cryptographic circuit is presented. Simulation results show that even though the power consumption is low, it can still be correlated with the correct key guess, and all possible subkeys applied to the Serpent sub-module can be revealed in a best case scenario for attackers. The security of practical QCA devices is also discussed and could be greatly improved by applying a smoother clock. Weiqiang Liu 0001, Saket Srivastava, Máire O'Neill, Earl E. Swartzlander Jr. |
ISCAS | 4 |
| 2013 | Partial encryption by randomized zig-zag scanning for video encodingabstractIn this paper, a novel partial encryption method for video encoding is proposed. For video compression based on block transforms, two zig-zag scan orders can be used which provide similar compression performance. The proposed method involves randomly selecting one of these two scan orders. In addition, the sign bit of the DC coefficients can be flipped randomly and when employed with the proposed method provides a much better scrambling effect. In comparison to previous work in this area, experimental results based on H.264/AVC show that the proposed method can scramble the video with very minimal impact on the compression performance. Máire O'Neill, Fatih Kurugollu |
ISCAS | 2 |
| 2013 | QCA Systolic Array DesignabstractQuantum-dot Cellular Automata (QCA) technology is a promising potential alternative to CMOS technology. To explore the characteristics of QCA and suitable design methodologies, digital circuit design approaches have been investigated. Due to the inherent wire delay in QCA, pipelined architectures appear to be a particularly suitable design technique. Also, because of the pipeline nature of QCA technology, it is not suitable for a complicated control system design. Systolic arrays take advantage of pipelining, parallelism, and simple local control. Therefore, an investigation into these architectures in semiconductor QCA technology is provided in this paper. Two case studies, (a matrix multiplier and a Galois Field multiplier) are designed and analyzed based on both multilayer and coplanar crossings. The performance of these two types of interconnections are compared and it is found that even though coplanar crossings are currently more practical, they tend to occupy a larger design area and incur slightly more delay. A general semiconductor QCA systolic array design methodology is also proposed. It is found that by applying a systolic array structure in QCA design, significant benefits can be achieved particularly with large systolic arrays, even more so than when applied in CMOS-based technology. Weiqiang Liu 0001, Máire O'Neill, Earl E. Swartzlander Jr. |
IEEE Trans. Computers | 3 |
| 2013 | A Tunable Encryption Scheme and Analysis of Fast Selective Encryption for CAVLC and CABAC in H.264/AVCabstractRecently, two fast selective encryption methods for context-adaptive variable length coding and context-adaptive binary arithmetic coding in H.264/AVC were proposed by Shahid In this paper, it was demonstrated that these two methods are not as efficient as only encrypting the sign bits of nonzero coefficients. Experimental results showed that without encrypting the sign bits of nonzero coefficients, these two methods can not provide a perceptual scrambling effect. If a much stronger scrambling effect is required, intra prediction modes, and the sign bits of motion vectors can be encrypted together with the sign bits of nonzero coefficients. For practical applications, the required encryption scheme should be customized according to a user's specified requirement on the perceptual scrambling effect and the computational cost. Thus, a tunable encryption scheme combining these three methods is proposed for H.264/AVC. To simplify its implementation and reduce the computational cost, a simple control mechanism is proposed to adjust the control factors. Experimental results show that this scheme can provide different scrambling levels by adjusting three control factors with no or very little impact on the compression performance. The proposed scheme can run in real-time and its computational cost is minimal. The security of the proposed scheme is also discussed. It is secure against the replacement attack when all three control factors are set to one. Máire O'Neill, Fatih Kurugollu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Application-oriented SHA-256 hardware design for low-cost RFIDabstractCryptographic hash functions can be used to provide strong security and privacy for Radio Frequency Identification (RFID) systems. In this paper, two application-oriented optimized SHA-256 hardware designs for low-cost RFID are presented. These are implemented on UMC 0.13 µm CMOS standard cell technology. One of the proposed designs achieves the smallest area (8,394 gates) while the other achieves the lowest power consumption (2.86 µW) in comparison to previous SHA-256 designs and SHA-3 finalist candidates reported in the literature to date. Furthermore, if the designs are used in low-cost RFID applications in which the input being authenticated is less than 400 bits, the two designs can be further optimized to utilise just 6,125 gates and consume 2.3 µW of power. Xiaolin Cao, Máire O'Neill |
ISCAS | 2 |
| 2012 | Cost-efficient decimal adder design in Quantum-dot cellular automataabstractApplications that cannot tolerate the loss of accuracy that results from binary arithmetic demand hardware decimal arithmetic designs. Binary arithmetic in Quantum-dot cellular automata (QCA) technology has been extensively investigated in recent years. However, only limited attention has been paid to QCA decimal arithmetic. In this paper, two cost-efficient binary-coded decimal (BCD) adders are presented. One is based on the carry flow adder (CFA) using a conventional correction method. The other uses the carry look ahead (CLA) algorithm which is the first QCA CLA decimal adder proposed to date. Compared with previous designs, both decimal adders achieve better performance in terms of latency and overall cost. The proposed CFA-based BCD adder has the smallest area with the least number of cells. The proposed CLA-based BCD adder is the fastest with an increase in speed of over 60% when compared with the previous fastest decimal QCA adder. It also has the lowest overall cost with a reduction of over 90% when compared with the previous most cost-efficient design. Weiqiang Liu 0001, Máire O'Neill, Earl E. Swartzlander Jr. |
ISCAS | 3 |
| 2012 | Adaptive binary mask for privacy region protectionabstractPrivacy region protection in video surveillance systems is an active topic at present. In previous research, a binary mask mechanism has been developed to indicate the privacy region; however this incurs a significant bitrate overhead. In this paper, an adaptive binary mask is proposed to represent the privacy region. In a practical privacy region protection application, in which the privacy region typically occupies less than half of the overall frame and is rectangular or approximately rectangular, the proposed adaptive binary mask can effectively reduce the bitrate overhead. The proposed method can also be easily applied to the FMO mechanism of H.264/AVC, providing both error resilience and a lower bitrate overhead. Máire O'Neill, Fatih Kurugollu |
ISCAS | 2 |
| 2011 | Power Spectral Density Side Channel Attack Overlapping Window MethodabstractCryptographic algorithms have been designed to be computationally secure, however it has been shown that when they are implemented in hardware, that these devices leak side channel information that can be used to mount an attack that recovers the secret encryption key. In this paper an overlapping window power spectral density (PSD) side channel attack, targeting an FPGA device running the Advanced Encryption Standard is proposed. This improves upon previous research into PSD attacks by reducing the amount of pre-processing (effort) required. It is shown that the proposed overlapping window method requires less processing effort than that of using a sliding window approach, whilst overcoming the issues of sampling boundaries. The method is shown to be effective for both aligned and misaligned data sets and is therefore recommended as an improved approach in comparison with existing time domain based correlation attacks. Philip Hodgers, Keanhong Boey, Máire O'Neill |
DSD | 3 |
| 2011 | Design rules for Quantum-dot Cellular AutomataabstractAs a promising alternative to CMOS technology, QCA circuit design has been extensively studied in recent years. However, although a concrete set of design rules exist for integrated circuit design, little attention has been paid to the design rules necessary for efficient QCA circuit design. This paper compiles a set of important QCA design rules which include layout design rules, timing rules and some special rules for QCA technology to ensure QCA circuits function correctly and reliably. These rules will promote the development of practical and efficient QCA systems. A GF(2m) multiplier design is proposed as a case study to illustrate these design rules. Weiqiang Liu 0001, Máire O'Neill, Earl E. Swartzlander Jr. |
ISCAS | 3 |
| 2011 | A Forward Private Protocol based on PRNG and LPN for Low-cost RFID
Xiaolin Cao, Máire O'Neill |
SECRYPT | 2 |
| 2011 | A Private and Scalable Authentication for RFID Systems Using Reasonable StorageabstractIn recent years, numerous authentication protocols for radio frequency identification (RFID) systems have been proposed to protect privacy. Due to the hardware resource limitation on RFID tags, the majority of these protocols are based on symmetric-key ciphers. However, most of them lack scalability at the reader side as they require a linear or logarithmic search proportional to the number of tags in the system in order to authenticate a tag. Recently proposed scalable protocols require a constant-time search and utilize hash look-up tables. However, the storage requirement of the hash-table in these protocols is very large, therefore not practical. Moreover, previous proposals do not support dynamic resizing where the number of tags changes during the life time of the RFID system. In this paper, a new Re-Hash technique is presented to reduce the hash-table storage in constant-time authentication protocols. To the best of our knowledge, it offers the smallest storage cost in comparison to previous proposals. The Re-Hash technique is also further adapted to support the dynamic scalability for RFID systems. Xiaolin Cao, Máire O'Neill |
TrustCom | 2 |
| 2010 | FPGA Implementations of the Round Two SHA-3 CandidatesabstractThe second round of the NIST-run public competition is underway to find a new hash algorithm(s) for inclusion in the NIST Secure Hash Standard (SHA-3). This paper presents the full implementations of all of the second round candidates in hardware with all of their variants. In order to determine their computational efficiency, an important aspect in NIST's round two evaluation criteria, this paper gives an area/speed comparison of each design both with and without a hardware interface, thereby giving an overall impression of their performance in resource constrained and resource abundant environments. The implementation results are provided for a Virtex-5 FPGA device. The efficiency of the architectures for the hash functions are compared in terms of throughput per unit area. To the best of the authors' knowledge, this is the first work to date to present hardware designs which test for all message digest sizes (224, 256, 384, 512), and also the only work to include the padding as part of the hardware for the SHA-3 hash functions. Brian Baldwin, Andrew Byrne, Mark Hamilton, Neil Hanley, Máire O'Neill, William P. Marnane |
FPL | 6 |
| 2010 | Lightweight DPA resistant solution on FPGA to counteract power modelsabstractElectronic cryptographic devices can be attacked by monitoring physical characteristics released from their circuits, such as power consumption and electromagnetic emanation. These techniques are known as Side Channel Attacks (SCAs). Differential Power Analysis (DPA) is one of the most effective SCAs, which can reveal the secret key from the dependency between the power consumption of the device and the processed data. This paper proposes a DPA resistant solution for FPGA implementations of the Advanced Encryption Standard (AES), combines two countermeasures, a new random inversion technique and an improved random register renaming countermeasure. This is the first time that the latter countermeasure is implemented on FPGA. The proposed solution achieves a very lightweight design in comparison to the previous countermeasures reported in the literature. Yingxi Lu, Keanhong Boey, Philip Hodgers, Máire O'Neill |
FPT | 4 |
| 2010 | Evaluation of Random Delay Insertion against DPA on FPGAsabstractSide-channel attacks (SCA) threaten electronic cryptographic devices and can be carried out by monitoring the physical characteristics of security circuits. Differential Power Analysis (DPA) is one the most widely studied side-channel attacks. Numerous countermeasure techniques, such as Random Delay Insertion (RDI), have been proposed to reduce the risk of DPA attacks against cryptographic devices. The RDI technique was first proposed for microprocessors but it was shown to be unsuccessful when implemented on smartcards as it was vulnerable to a variant of the DPA attack known as the Sliding-Window DPA attack. Previous research by the authors investigated the use of the RDI countermeasure for Field Programmable Gate Array (FPGA) based cryptographic devices. A split-RDI technique was proposed to improve the security of the RDI countermeasure. A set of critical parameters was also proposed that could be utilized in the design stage to optimize a security algorithm design with RDI in terms of area, speed and power. The authors also showed that RDI is an efficient countermeasure technique on FPGA in comparison to other countermeasures. In this article, a new RDI logic design is proposed that can be used to cost-efficiently implement RDI on FPGA devices. Sliding-Window DPA and realignment attacks, which were shown to be effective against RDI implemented on smartcard devices, are performed on the improved RDI FPGA implementation. We demonstrate that these attacks are unsuccessful and we also propose a realignment technique that can be used to demonstrate the weakness of RDI implementations. Yingxi Lu, Máire O'Neill, John V. McCanny |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2009 | Common Control Channel Security Framework for Cognitive Radio NetworksabstractCognitive radio networks are becoming an increasingly important part of the wireless networking landscape due to the ever-increasing scarcity of spectrum resources. Such networks perform co-operative spectrum sensing to find white spaces and apply policies to determine when and in which bands they may communicate. In a typical MAC protocol designed for cooperatively communicating ad hoc cognitive radio networks, nodes make use of a common control channel to perform channel negotiations before any actual data transmission. The provision of common control channel security is vital to ensure any subsequent security among the communicating cognitive radio nodes. To date, wireless security has received little attention in cognitive radio networks research. The cognitive radio paradigm introduces entirely new classes of security threats and challenges, such as selfish misbehaviours, licensed user emulation and eavesdropping. This paper presents a novel framework for providing common control channel security for co-operatively communicating cognitive radio nodes. To the best of the authors' knowledge, this is the first paper which proposes such a concept. The paper investigates how two cognitive radio nodes can authenticate each other prior to any confidential channel negotiations to ensure subsequent security against attacks. The paper also describes the importance of common control channel security and concludes with future work describing the realization of the proposed framework. Ghazanfar Ali Safdar, Máire O'Neill |
VTC Spring | 2 |
| 2008 | FPGA implementation and analysis of random delay insertion countermeasure against DPAabstractSecurity devices can reveal critical information about the cryptographic key from the power consumption of their circuits. Differential Power Analysis (DPA) is one of the most effective power analysis techniques. In recent years numerous countermeasures against the DPA attack of hardware implementations of security algorithms have been proposed. In this paper, we investigate the Random Delay Insertion (RDI) countermeasure. Previous research has evaluated RDI for microprocessor implementations; however, its security properties in relation to hardware implementations have not been investigated in detail. We prove both theoretically and practically that it is an effective technique on FPGA devices and we propose a set of critical parameters that can be utilized to optimize a security algorithm design with RDI in terms of area, speed and power. In this work, we implement the first hardware security architecture with RDI on an FPGA device, and attack it using DPA. It is shown that RDI is an efficient countermeasure technique on FPGA in comparison to other countermeasures. Yingxi Lu, Máire O'Neill, John V. McCanny |
FPT | 2 |
| 2008 | Differential Power Analysis of a SHACAL-2 hardware implementationabstractSide-channel attacks (SCA) can be used to reveal the security key stored in cryptographic implementations by monitoring characteristics such as power consumption and subsequently applying statistical analysis techniques. However, cryptographic algorithms, such as SHACAL-2, which do not have any S-Box operations, are more resistant to SCA than algorithms that do contain S-Box operations, such as the S-box computation in DES or AES. In this paper, the advantages of SHACAL-2 in relation to its resistance to Differential Power Analysis (DPA) are outlined. The first reported DPA attack of a SHACAL-2 encryption algorithm FPGA implementation is also presented. Finally, the effect of using different power models in the DPA attack of SHACAL-2 is discussed. Yingxi Lu, Máire O'Neill, John V. McCanny |
ISCAS | 2 |
| 2007 | Public Key Cryptography and RFID Tags
Máire O'Neill, Matthew J. B. Robshaw |
CT-RSA | 1 |
| 2007 | New Architectures for Low-Cost Public Key Cryptography on RFID TagsabstractAlthough it is commonly believed that the computational complexity of public key cryptography prevents its deployment on low-cost RFID tags, it was recently demonstrated (McLoone and Robshaw, 2007) that the GPS identification scheme provides a counter-example to this view; with regards to all three attributes of space, power, and timing, GPS is well-suited to low-cost implementation. In this paper we consider new and innovative hardware architectures for implementing the GPS identification scheme and these allow a broader range of practical performance trade-offs. Máire O'Neill, Matthew J. B. Robshaw |
ISCAS | 1 |
| 2007 | Identity Based Public Key Exchange (IDPKE) for Wireless Ad Hoc Networks
Clare McGrath, Ghazanfar Ali Safdar, Máire O'Neill |
SECRYPT | 3 |
| 2007 | MONET Special Issue on Next Generation Hardware Architectures for Secure Mobile Computing
Nicolas Sklavos 0001, Máire O'Neill, Xinmiao Zhang 0001 |
Mob. Networks Appl. | 2 |
| 2006 | An Adaptable And Scalable Asymmetric Cryptographic ProcessorabstractIn this paper a novel scalable public-key processor architecture is presented that supports modular exponentiation and Elliptic Curve Cryptography over both prime GF(p) and binary GF(2n) extension fields. This is achieved by a high performance instruction set that provides a comprehensive range of integer and polynomial basis field arithmetic. The instruction set and associated hardware are generic in nature and do not specifically support any cryptographic algorithms or protocols. Firmware within the device is used to efficiently implement complex and data intensive arithmetic. A firmware library has been developed in order to demonstrate support for numerous exponentiation and ECC approaches, such as different coordinate systems and integer recoding methods. The processor has been developed as a high-performance asymmetric cryptography platform in the form of a scalable Verilog RTL core. Various features of the processor may be scaled, such as the pipeline width and local memory subsystem, in order to suit area, speed and power requirements. The processor is evaluated and compares favourably with previous work in terms of performance while offering an unparalleled degree of flexibility. Neil Smyth, Máire O'Neill, John V. McCanny |
ASAP | 2 |
| 2005 | High-Radix Systolic Modular Multiplication on Reconfigurable Hardware
Ciaran McIvor, Máire O'Neill, John V. McCanny |
FPT | 2 |
| 2005 | High-Speed Hardware Architectures of the Whirlpool Hash Function
Máire O'Neill, Ciaran McIvor, Aidan Savage |
FPT | 1 |
| 2004 | FPGA Montgomery Multiplier Architectures - A ComparisonabstractNovel FPGA architectures for the SOS, CIOS and FIOS Montgomery multiplication algorithms are presented. The 18/spl times/18-bit multipliers and fast carry look-ahead logic embedded within the Xilinx Virtex2 Pro family of FPGAs are used to perform the ordinary multiplications and additions required by these algorithms. A detailed analysis is given, highlighting the advantages and weaknesses of each of these architectures when implemented in hardware. This shows that the CIOS multiplier architectures perform best overall, with the performance gap between this and the other options increasing as the word size used decreases. In addition, the SOS multipliers outperform the FIOS multipliers for larger word sizes, but vice versa as the word size decreases. It is also shown that one can tailor the multiplier architectures to be area efficient, time efficient or a mixture of both, by choosing a particular word size. Ciaran McIvor, Máire O'Neill, John V. McCanny |
FCCM | 2 |
| 2004 | Coarsely integrated operand scanning (CIOS) architecture for high-speed Montgomery modular multiplicationabstractA generic coarsely integrated operand scanning (CIOS) architecture that provides high speed Montgomery modular multiplication is presented in This work. The architecture is capable of supporting varying operand sizes. It achieves a throughput of 210 Mbps, 289 Mbps and 334 Mbps for 128-bit, 256-bit and 512-bit operand sizes respectively, when implemented on a Virtex XC2 VP50 FPGA. Throughputs of up to 400 Mbps are achieved if the final subtraction in the Montgomery algorithm is excluded. To the authors' knowledge this is the fastest Montgomery multiplication architecture reported in the literature. Máire O'Neill, Ciaran McIvor, John V. McCanny |
FPT | 1 |
| 2003 | Very High Speed 17 Gbps SHACAL Encryption Architecture
Máire O'Neill, John V. McCanny |
FPL | 1 |
| 2002 | Efficient single-chip implementation of SHA-384 and SHA-512abstractThe rapid developments in the communications industry over the last decade have led to an escalation in the amount of sensitive data being transmitted over the Internet. This has resulted in an increased awareness of the need to provide security measures. Authentication is one such security measure. A novel highly efficient single-chip hardware design of the SHA-384 and SHA-512 authentication algorithms is described in this paper. The compact implementation achieves a throughput of 479 Mbits/sec utilising a shift register design approach and look-up tables (LUTs). This is believed to be the first SHA-384/SHA-512 hardware implementation to be reported in the literature. Máire O'Neill, John V. McCanny |
FPT | 1 |
| 2001 | High Performance Single-Chip FPGA Rijndael Algorithm Implementations
Máire O'Neill, John V. McCanny |
CHES | 1 |
| 2001 | Single-Chip FPGA Implementation of the Advanced Encryption Standard Algorithm
Máire O'Neill, John V. McCanny |
FPL | 1 |