EDBT 2026 Demo / reviewers in the wild / expert
Ayesha Khalid
dblp:33/10570
· DBLP profile ↗
33ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 5 first-author · 13 since 2021Security and privacy · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Planting CRYSTALS-Kyber Acceleration SeedsabstractCRYSTALS-Kyber (FIPS203) was one of the first post quantum cryptography schemes to be standardized by the National Institute of Standards and Technology; CRYSTALS-Kyber utilizes Montgomery reduction at the heart of the modulo reduction. Montgomery reduction traditionally is suited to computing moduli much larger than 32/64-bits. Recently a new modulo reduction method was published named Plantard reduction which allows for one less multiplication at the reduction stage. However, this does require doubling the word size of the values being calculated, in the case of CRYSTALS-Kyber increasing from 16-bit to 32-bit. In this paper we present the worlds first Plantard ISE on the Ibex RISC-V core and investigate the impact Instruction Set Extensions can have utilizing this new modulo reduction method over original methods such as Montgomery and Barrett reduction. Utilizing Plantard reduction as a set of Instruction Set Extensions, we are able to reduce the cycle counts required to compute the entire CRYSTALS-Kyber algorithm by up to 21% and reduce the cycle counts of the NTT by 1.7x and the INTT by 2.6x against the reference implementation. Ryan Bevin, Ayesha Khalid, Seog Chung Seo, Máire O'Neill |
ISCAS | 2 |
| 2026 | HBQS: Lightweight Post-Quantum Secure Authentication for Satellite Networks Leveraging Hardware TRNG and PUFsabstractSatellite communication networks play a critical role in providing connectivity to remote regions and areas with limited infrastructure. However, their inherently open nature and physical exposure make them particularly susceptible to security threats, including replay, impersonation, and man-in- the-middle attacks. The emergence of quantum computing further undermines the robustness of conventional cryptographic schemes that rely on number-theoretic assumptions. To mitigate these challenges, this article proposesHBQS, a lightweight post-quantum authentication framework designed for satellite platforms with limited resources.HBQSintegrates Physically Unclonable Functions (PUFs) with hash-based cryptography, leveraging SPHINCS+ digital signatures and SHA-3 hashing to provide secure mutual authentication and session key establishment. The protocolHBQSwas implemented and evaluated on embedded hardware platforms, including Raspberry Pi 4.0, PYNQ-Z2 FPGA, and a Dell ground control station. The experimental results demonstrate thatHBQSachieves mutual authentication in 0.88 ms, offering an approximately 67% reduction in execution time compared to representative baselines from the prior literature. Entropy analysis confirms that the proposed protocol maintains a high entropy across critical components, withHBQSachieving a signature entropy of 164.63 bits and PUF response entropy of 172.58 bits, indicating strong resistance to statistical and modeling attacks. A formal security analysis conducted within the Random Oracle Model demonstrates semantic security against both classical and quantum adversaries. The protocolHBQSshows a significant improvement over existing methods, achieving approximately 67% reduction in computational execution time, approximately 11.5% faster user-side handshake time, approximately 0.6% improvement on the satellite side and full security coverage across all evaluated security features with only a 21.4% increase in static memory usage as the trade-off. These findings positionHBQSas an efficient, secure, and scalable authentication solution for next-generation satellite communication systems operating in the post-quantum era. Muhammad Arslan Akram, Arnab Kumar Biswas, Máire O'Neill, Ayesha Khalid, Adnan Noor Mian |
IEEE Internet Things J. | 4 |
| 2026 | LightHD: A Lightweight and High-Performance Hardware Accelerator of CRYSTALS-DilithiumabstractCRYSTALS-Dilithium serves as the foundation of the NIST-standardised PQC digital signature scheme, and has been declared as the first recommended digital signature algorithm. However, due to the computational complexity and intricate processing flow of CRYSTALS-Dilithium, two limitations are shown in existing methods: its applicability on resource-constrained devices is limited and the performance reported so far remains relatively low. This paper presents a lightweight yet high-performance hardware architecture that optimises the core computational units of CRYSTALS-Dilithium. First, an iterative dual-Keccak SHA-3 module is proposed, where two cores operate with a 26-cycle offset to accelerate processing without compromising frequency. In addition, the rejection sampler is streamlined by two compact registers for intermediate values and counters, improving efficiency when consuming interleaved SHA-3 outputs. Second, for small bit-width polynomials, we eliminate the first NTT stage via lookup tables and data regrouping, reducing NTT cycles by 11.7% with little hardware overhead. Further hardware savings are achieved by maximising IP core utilisation and simplifying input multiplexers. Furthermore, a compact scheduling strategy ensures that all intermediate storage fits within a single polynomial-sized memory block. On Xilinx Artix-7 FPGAs, the design reduces hardware overhead by 14.2% compared with state-of-the-art lightweight implementations. Across three security levels, KeyGen and Verify are 27.3% and 13.5% faster, respectively, than high-performance prior designs. At level 5, the best-case Sign latency is only 120 µs. Ziying Ni, Ayesha Khalid, Zhaoyu Zhang 0001, Yijun Cui, Weiqiang Liu 0001, Máire O'Neill |
IEEE Trans. Computers | 2 |
| 2025 | Kyber-KEM-Ascon: Benchmarking a Lightweight Post-quantum KEM on IoT DevicesabstractCRYSTALS-Kyber, officially standardized by the U.S. National Institute of Standards and Technology (NIST) in August 2024 as ML-KEM FIPS, is the only post quantum secure key encapsulation mechanism (KEM) standard. Due to its inherent computationally intensive nature, mapping it on lightweight IoT devices is often a struggle. This study examines the replacement of the Keccak hashing function in Kyber-KEM with the NIST LWC competition winner called Ascon hash function. We compare speed/ memory improvement by executing Kyber-KEM-Ascon on an ARM Cortex M4 device and present novel benchmarking results; Kyber-KEM-Ascon shows about 25% reduction in clock cycles, along with lower memory usage requirements (about 8%). These results suggest that Kyber-KEM-Ascon is more suited for lightweight platforms, offering benefits for IoT security applications. Sahar Shehzadi, Nathan Whaley, Ayesha Khalid, Abdul Ghafoor 0002, Sadiqa Arshad, Faiz Ul Islam, Máire O'Neill |
ISCAS | 3 |
| 2025 | An Enhanced Two-Step CPA Side-Channel Analysis Attack on ML-KEMabstractThis work presents an enhanced two-step Correlation Power Analysis (CPA) attack targeting the recently standardised ML-KEM on an ARM Cortex M4. Our enhancement exploits the knowledge of intermittent variables to identify sample points of interest and develop bespoke attack functions. Step one targets the odd coefficients of each Secret Key Polynomial Vector ( ˆ s), before step two targets the remaining even coefficients using more elaborate attack functions. After successfully demonstrating key recovery for the first set of ˆ s, we then characterise leakage behaviour, revealing a trend indicating recovery of each coefficient becomes more efficient with subsequent iterations of the internal doublebasemul operation. By applying our enhanced two step attack methodology, we successfully recovered the entire key using only 179 traces, without the need for elaborate preconditions or ciphertext manipulations. We obtain remarkable results in the initial stage of our attack, while the second phase achieves performance comparable to other recent studies. Mark Kennaway, Anh-Tuan Hoang, Ayesha Khalid, Ciara Rafferty, Máire O'Neill |
SECRYPT | 3 |
| 2025 | A SWOT Analysis of Software Development Life Cycle Security MetricsabstractABSTRACT Cyber security is an ongoing and critical concern due to persistent threats posed by threat actors, such as hackers and crackers. With the development of information and communication technologies (ICT), the widespread usage of software systems has transformed modern society in many ways but also created new issues in protecting confidential and sensitive information. The quantification of security measures can provide evidence to support decision‐making in software security, particularly when assessing the security performance of software systems. This entails understanding the key quality criteria of security metrics, which can assist in constructing security models aligned with practical requirements. To delve deeper into this subject, the current study conducted a systematic literature review (SLR) on security metrics and measures within the realm of secure software development (SSD). The study selected 61 research publications for data extraction based on the specific inclusion and exclusion criteria. The study identified 215 software security metrics and classified them into different phases of software development life cycle (SDLC). In order to evaluate the most cited metrics in each phase of SDLC, the strengths, weaknesses, opportunities, and threats (SWOT) analysis was performed. The SWOT analysis offers a structured framework enabling researchers to make more effective, well‐informed decisions and mitigate potential risks, ultimately contributing to more valuable research findings. The study's findings provide researchers guidance for exploring emerging trends and addressing existing gaps in SDLC. This study also provides software professionals with a more comprehensive understanding of security measurements, constraints, and open‐ended specific and general issues. Ayesha Khalid, Mushtaq Raza, Palwasha Afsar, Rafiq Ahmad Khan, Muhammad Ismail Mohmand, Hanif Ur Rahman |
J. Softw. Evol. Process. | 1 |
| 2025 | A Highly Hardware Efficient ML-KEM Accelerator with Optimised Architectural LayersabstractThe Module-Lattice-Based Key encapsulation Mechanism (ML-KEM) scheme, which is currently being standardised, is a quantum attack resistant KEM that is based on CRYSTALS-Kyber. CRYSTALS-Kyber is the only Public-key Encryption (PKE)/ KEM scheme selected in the first set of successful candidates as part of the NIST initiated Post-Quantum Cryptography (PQC) process. ML-KEM scheme includes three different security levels, namely security level 1, 3, and 5. In this research, we propose a highly area-time efficient hardware ML-KEM architecture. The architecture comprises three computational layers. The first layer comprises a hash and sampling module; the second layer includes a number theoretic transform (NTT), its inverse (INTT) and a point-wise multiplication (PWM) module; and the third layer comprises addition, compressing and encoding. Intra-layer pipelining and out-of-layer scheduling ensures that either layer 1 or layer 2 operate in the shortest time. In the reduction module, we propose a novel hybrid architecture to obtain the final result within 2 cycles with low area consumption. In the NTT module, the PWM pipelining method is modified and an optimised iterative FIFO access method is adopted to reduce the size of FIFO units by 55% over previous research. Look-up tables are also used to replace the first-stage of the NTT to reduce 8 cycles. Furthermore, the memory unit uses only FIFOs, the size are optimised based on the requirements of the most resource-intensive function in ML-KEM (ML-KEM.CPA.Dec). The results show that the proposed architecture has a 48.2%, 41.2%, and 78.1% reduction in computational time in comparison to previous work for security levels 1, 3, and 5, respectively. In addition, the area of proposed optimised ML-KEM designs is reduced by 73%, 70%, 76% and resulting in an improved area-time (AT) product of 15.8%, 10.7%, and 11.3%, for the Level 1, 3, and 5 security levels respectively, compared with state-of-the-art designs. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | Bitstream Fault Injection Attacks on CRYSTALS Kyber Implementations on FPGAsabstractCRYSTALS-Kyber is the only Public-key Encryption (PKE)/ Key-encapsulation Mechanism (KEM) scheme that was chosen for standardization by the National Institute of Standards and Technology initiated Post-quantum Cryptography competition (so called NIST PQC). In this paper, we show the first successfully malicious modifications of the bitstream of a Kyber FPGA implementation. We successfully demonstrate 4 different attacks on Kyber hardware implementations on Artix-7 FPGAs that either reduce the complexity of polynomial multiplication operations or enable direct secret key/ message recovery by: disabling BRAMs, disabling DSPs, zeroing NTT ROM and tampering with CBD2 results. Two of our attacks are generic in nature and the other two require reverse-engineering or a detailed knowledge of the design. We evaluate the feasibility of the four attacks, among which the zeroing NTT ROM and tampering with the CBD2 result attacks produce higher public key and ciphertext complexity and thus are difficult to be detected. Two countermeasures are proposed to prevent the attacks proposed in this paper. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
DATE | 2 |
| 2024 | FPGA Bitstream Fault Injection Attack and Countermeasures on the Sampling Counter in CRYSTALS KyberabstractThe CRYSTALS Kyber algorithm is the public key encryption (PKE)/ key encapsulation mechanism (KEM) protocol undertaken for standardization by the US National Institute of Standards and Technology (NIST) the PQC competition and serves as the foundation for the Module-Lattice-Based (ML)-KEM scheme. The inherently strong security properties of the Kyber algorithm are considered to be resistant to attacks under quantum computers, but the security of its FPGA-based hardware implementation circuitry is still worth considering. In this work, we introduce the Nonce counter disabling attack, which targets the binomial distribution sampling process. We demonstrate that, in the modified primes of Kyber from Round 2, it is also effectively deduce the secret key s by equating it with the noise e. Our implementation of this attack on a Nexys 4 FPGA, with an additional DSP disabling filtering process to pinpoint the LUT. This attack is applicable to both the key generation and key encapsulation phases, and only need to modify 32-bit bitstream. Finally, We propose the Nonce counter check and the splitting of the Nonce computation cycles methods to to prevent this attack in hardware design-level. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 2 |
| 2024 | Efficient Soft Core Multiplier for Post Quantum Digital SignaturesabstractMultiplication is a core operation in various applications such as cryptography and machine learning. Dedicated DSP blocks are provided by FPGA vendors for multiplication. However, these DSP blocks are limited in number and their location on FPGA is fixed, resulting in routing delays that affects the performance for small size multipliers. In this paper, a high performance and resource efficient 5 × 5 multiplier is presented that utilizes lookup tables (LUTs) and fast carry chain of the FPGA. The proposed multiplier offers 30% reduction in LUTs compared to Vivado DSP-less inferred multiplier at the cost of a slight increase in critical path delay (CPD). The proposed multiplier requires lesser power consumption and has better area- delay product (ADP) and power-delay product (PDP) metrics. Based on the proposed multiplier, a finite field multiplier is developed for post quantum digital signatures such as QR-UOV, MAYO and MQOM. The matrix-vector architecture is the core operation in multivariate digital signatures and integration of our finite field multiplier in a matrix-vector architecture shows that area is almost halved compared to state-of-the-art. Yasir Ali Shah, Ciara Rafferty, Ayesha Khalid, Safiullah Khan, Khalid Javeed, Máire O'Neill |
ISCAS | 3 |
| 2023 | Towards a Lightweight CRYSTALS-Kyber in FPGAs: an Ultra-lightweight BRAM-free NTT CoreabstractCRYSTALS-Kyber is the first quantum-resilient, lattice-based Public Key Encryption (PKE)/Key Encapsulation Mechanism (KEM) cryptosystem that is chosen by the ongoing National Institute of Standards and Technology post-quantum cryptography standardization (NIST PQC) for standardization. This work presents a lightweight and efficient, FPGA-based hardware implementation for polynomial multiplication unit (NTT), which is the major bottleneck in the Kyber scheme. As a first step, an optimzed modular multiplication architecture combining KRED and lookup table-based algorithms is presented, which reduces the resources of slices by 16.7%. It is used in a pipelined NTT/INTT architecture that is completely BRAM free and instead uses 3 FIFOs for coefficients storage. We hereby present the most compact FPGA based design for NTT architecture in Kyber till date. Experimental results bench marked on comparable FPGA devices show that our proposed design is 36-75% better than the state-of-the-art implementations in terms of hardware efficiency for NTT/INTT calculations and$3.4-4.4\times$better for the Point-wise Multiplication (PWM) operation. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 2 |
| 2023 | Efficient, Error-Resistant NTT Architectures for CRYSTALS-Kyber FPGA AcceleratorsabstractThe dawn of cost-effective miniaturised satellites is currently attracting venture capital in a never seen before ratio to launch mega-constellations of satellites for a diverse range of applications. These satellites are vulnerable to attacks by high-capability cyber-criminals (including quantum enabled adversaries), due to the critical data they transmit. Additionally, space missions have long lifespan and a long lead time in terms of development process, requiring a pre-emptive outlook to ensuring their safety. In 2016, National Institute of Standards and Technology (NIST) initiated the competition to standardise the post-quantum cryptography (PQC) schemes, announcing the first portfolio of chosen schemes in 2022. This work targets the only public key exchange (PKE) scheme among the winners of the NIST-PQC standardisation process, CRYSTALS-Kyber, and implements its core bottleneck operation, i.e., number theoretic transform (NTT) extensively used for the polynomial multiplication. To avoid data corruption due to space based radiations, a novel error-resistant model for NTT is presented based on hybrid protection mechanisms, i.e., the use of hamming codes for detection and correction of errors in the twiddle factors and the use of parity computed for all NTT coefficients for error detection. Benchmarking error protection overheads on a Xilinx Virtex-7 FPGA reports 16.4% and 10.8% degradation on the hardware efficiency when the hamming codes for twiddle factors and parity bit for NTT coefficients are used to mitigate errors, respectively. A total of 29.2% area overhead is benchmarked when compared to the standard unprotected NTT implementations. Safiullah Khan, Ayesha Khalid, Ciara Rafferty, Yasir Ali Shah, Máire O'Neill, Wai-Kong Lee, Seong Oun Hwang |
VLSI-SoC | 2 |
| 2023 | HPKA: A High-Performance CRYSTALS-Kyber Accelerator Exploring Efficient PipeliningabstractCRYSTALS-Kyber (Kyber) was recently chosen as the first quantum resistant Key Encapsulation Mechanism (KEM) scheme for standardisation, after three rounds of the National Institute of Standards and Technology (NIST) initiated PQC competition which begin in 2016 and search of the best quantum resistant KEMs and digital signatures. Kyber is based on the Module-Learning with Errors (M-LWE) class of Lattice-based Cryptography, that is known to manifest efficiently on FPGAs. This work explores several architectural optimizations and proposes a high-performance and area-time (AT) product efficient hardware accelerator for Kyber. The proposed architectural optimizations include inter-module and intra-module pipelining, that are designed and balanced via FIFO based buffering to ensure maximum parallelisation. The implementation results show that compared to state-of-the-art designs, the proposed architecture delivers 25–51% speedups for Kyber's three different security levels on Artix-7 and Zynq UltraScale+ devices, and a 50–75% reduction in DSPs at comparable security level. Consequently, the proposed design achieve higher AT product efficiencies of 19–33%. Ziying Ni, Ayesha Khalid, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Computers | 2 |
| 2023 | KaratSaber: New Speed Records for Saber Polynomial Multiplication Using Efficient Karatsuba FPGA ArchitectureabstractSABER is a round 3 candidate in the NIST Post-Quantum Cryptography Standardization process. Polynomial convolution is one of the most computationally intensive operation in Saber Key Encapsulation Mechanism, that can be performed through widely explored algorithms like the schoolbook polynomial multiplication algorithm (SPMA) and Number Theoretic Transform (NTT). While SPMA multiplier has a slow latency performance, the NTT-based multiplier usually requires large hardware. In this work, we propose KaratSaber, an optimized Karatsuba polynomial multiplier architecture with a balanced hardware efficiency (throughput-per-slice, TPS) compared to NTT and SPMA based designs. KaratSaber employs several techniques for an efficient design: a parallel grid input technique for efficient pre-processing stage in Karatsuba-based polynomial multiplier, a novel instruction code result-mapping technique catering the negacyclic operations improves the post-processing stage efficiency, a double multiplicand shifter-based multiplier doubles the throughput at the multiplication stage. Combining these three techniques, the proposed KaratSaber architecture is 7.47 × faster compared to the state-of-the-art SPMA Saber architecture at the expense of 4.96 × additional hardware resources; making KaratSaber 46.04% more area-time efficient. When compared to LWRPro, a recent Karatsuba Saber architecture, KaratSaber architecture achieves a 2.11 × higher throughput by only utilizing 1.92 × additional hardware; thus gaining a 10.44% improvement in area-time efficiency. Zheng-Yan Wong, Denis Chee-Keong Wong, Wai-Kong Lee, Kai Ming Mok, Wun-She Yap, Ayesha Khalid |
IEEE Trans. Computers | 6 |
| 2022 | Acceleration of Post Quantum Digital Signature Scheme CRYSTALS-Dilithium on Reconfigurable HardwareabstractThis research investigates efficient architectures for the implementation of the CRYSTALS-Dilithium post-quantum digital signature scheme on reconfigurable hardware, in terms of speed, memory usage, power consumption and resource utilisation. Post quantum digital signature schemes involve a significant computational effort, making efficient hardware accelerators an important contributor to future adoption of schemes. This is work in progress, comprising the establishment of a comprehensive test environment for operational profiling, and the investigation of the use of novel architectures to achieve optimal performance. Donal Campbell, Ciara Rafferty, Ayesha Khalid, Máire O'Neill |
FPL | 3 |
| 2022 | High Performance FPGA-based Post Quantum Cryptography ImplementationsabstractPost-quantum Cryptography (PQC) is an umbrella term for cryptographic schemes based on hard mathematical problems which are resistant to attacks by quantum computers. The National Institute of Standards and Technology (NIST) initiated a PQC standardisation process in 2017, with a total of 4 algorithms selected for standardisation after round 3 and 4 undertaken for further analysis in Round 4 in 2022. PQC schemes on hardware devices, such as Field Programmable Gate Arrays (FPGA), show the potential of higher throughput performance, for comparable security, at the cost of high area and power consumption. The major aim of this thesis is to help facilitate the global transition to a post quantum secure set of security protocols. This thesis will focus on the optimisation of the the hardware architectures to improve the computational speed and reduce the area overhead. The side channel analysis vulnerabilities and their countermeasures will also be studied. Ziying Ni, Ayesha Khalid, Máire O'Neill |
FPL | 2 |
| 2022 | Stacked Ensemble Model for Enhancing the DL based SCAabstractDeep learning (DL) has proven to be very effective for image recognition tasks, with a large body of research on various models for object classification. The application of DL to side-channel analysis (SCA) has already shown promising results, with experimentation on open-source variable key datasets showing that secret keys for block ciphers like Advanced Encryption Standard (AES)-128 can be revealed with 40 traces even in the presence of countermeasures. This paper aims to further improve the application of DL in SCA, by enhancing the power of DL when targeting the secret key of cryptographic algorithms when protected with SCA countermeasures. We propose a stacked ensemble model, which trains the output probabilities and Maximum likelihood score of multiple traces and/or sub-models to improve the performance of Convolutional Neural Network (CNN)-based models. Our model generates state-of-the art results when attacking the ASCAD variable-key database, which has a restricted number of training traces per key, recovering the key within 20 attack traces in comparison to 40 traces as required by the state-of-the-art CNN-based model with Plaintext feature extension (CNNP)-based model. During the profiling stage an attacker needs no additional knowledge of the implementation, such as the masking scheme or random mask values, only the ability to record the power consumption or electromagnetic field traces, plaintext/ciphertext and the key is needed. However, a two step training procedure is required. Additionally, no heuristic pre-processing is required in order to break the multiple masking countermeasures of the target implementation. Anh-Tuan Hoang, Neil Hanley, Ayesha Khalid, Dur-e-Shahwar Kundi, Máire O'Neill |
SECRYPT | 3 |
| 2022 | AxRLWE: A Multilevel Approximate Ring-LWE Co-Processor for Lightweight IoT ApplicationsabstractThis work presents a multilevel approximation exploration undertaken on the Ring-Learning-with-Errors (R-LWE)-based public-key cryptographic (PKC) schemes that belong to quantum-resilient cryptography algorithms. Among the various quantum-resilient cryptography schemes proposed in the currently running NIST’s post-quantum cryptography (PQC) standardization plan, the lattice-based learning-with-error (LWE) schemes have emerged as the most viable and preferred class for the Internet of Things (IoT) applications due to their compact area and memory footprint compared to other alternatives. However, compared to the classical schemes used today, R-LWE is much harder a challenge to fit on embedded IoT (end-node) devices, due to their stricter resource constraints (lower area, memory, and energy budgets) as well as their limited computational capabilities. To the best of our knowledge, this is the first endeavor exploring the inherent approximate nature of the LWE problem to undertake a multilevel approximate R-LWE (AxRLWE) architecture with respective security estimates opt for lightweight IoT devices. Undertaking AxRLWE on field-programmable gate arrays (FPGAs), we benchmarked a 64% area reduction cost compared to earlier accurate R-LWE designs at the cost of reduced quantum security. For the application-specific integrated circuits (ASICs) with 45-nm CMOS technology, AxRLWE was benchmarked to fit well within the same area budget of a lightweight ECC processor and consume a third of energy compared to special class of R-Binary LWE (R-BLWE) designs being proposed for an IoT, with a better security level. Dur-e-Shahwar Kundi, Ayesha Khalid, Song Bian 0001, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Internet Things J. | 2 |
| 2020 | A Secure Algorithm for Rounded Gaussian Sampling
Séamus Brannigan, Máire O'Neill, Ayesha Khalid, Ciara Rafferty |
CANS | 3 |
| 2020 | AxMM: Area and Power Efficient Approximate Modular Multiplier for R-LWE CryptosystemabstractAmongst various Post-Quantum Cryptographic (PQC) schemes, Lattice-Based Cryptography (LBC) stands out as the most viable substitute to the classical cryptographic schemes due to its efficiency, versatility and solid foundations on hard mathematical problems. Ring Learning With Errors (R-LWE) is a Public Key Encryption (PKE) scheme of LBC, in which the modular polynomial multiplication in a ring is the main bottleneck in the realization of a practical resource-constraint design for the embedded IoT devices. This work explores novel Approximate Computing (AC) technique for the design of area/power efficient modular multiplier (so called AxMM) for R-LWE, exploiting the inherent approximate structure of the scheme. The proposed AxMM on 45nm ASIC library achieved an area and power reduction of 36% and 23%, respectively, along with a speed increase of 1.34× as compared to state-of-art smallest exact R-LWE modular multiplier. Dur-e-Shahwar Kundi, Song Bian 0001, Ayesha Khalid, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2019 | Fault Attack Countermeasures for Error Samplers in Lattice-Based CryptographyabstractLattice-based cryptography is one of the leading candidates for NIST's post-quantum standardisation effort, providing efficient key encapsulation and signature schemes. Most of these schemes base their hardness on variants of LWE, and thus rely heavily on error samplers to provide necessary uncertainty by obfuscating computations on secret information. Because of this it is a clear and obvious target for side-channel analysis, with numerous types of attacks targeting this component to gain secret-key information. In order to bring potential lattice-based cryptographic standards to practical realisation, it is important to protect these modules from past and future fault and side-channel attacks. This paper proposes countermeasures that exploit the distributions expected from these error samples, that is either Gaussian or binomial, by using statistical tests to verify the samplers are operating properly. The novel countermeasures are designed to protect against all previous fault attacks on error samplers. We optimize hardware implementation of the proposed tests to avoid division and square root calculations, however, the countermeasure we propose is sufficiently generic to be suitable also for software. We measure the impact of these countermeasures on performance and area consumption on a Xilinx Artix-7 FPGA. Our countermeasure achieve promising performance while resulting in a minimal overhead. James Howe, Ayesha Khalid, Marco Martinoli, Francesco Regazzoni 0001, Elisabeth Oswald |
ISCAS | 2 |
| 2019 | Optimized Schoolbook Polynomial Multiplication for Compact Lattice-Based Cryptography on FPGAabstractLattice-based cryptography (LBC) is one of the most promising classes of post-quantum cryptography (PQC) that is being considered for standardization. This brief proposes an optimized schoolbook polynomial multiplication (SPM) for compact LBC. We exploit the symmetric nature of Gaussian noise for bit reduction. Additionally, a single field-programmable gate array (FPGA) DSP block is used for two parallel multiplication operations per clock cycle. These optimizations enable a significant 2.2× speedup along with reduced resources for dimension n = 256. The overall efficiency (throughput per slice) is 1.28× higher than the conventional SPM, as well as contributing to a more compact LBC system compared to previously reported designs. The results targeting the FPGA platform show that the proposed design can achieve high hardware efficiency with reduced hardware area costs. Weiqiang Liu 0001, Sailong Fan, Ayesha Khalid, Ciara Rafferty, Máire O'Neill |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Physical Protection of Lattice-Based Cryptography: Challenges and SolutionsabstractThe impending realization of scalable quantum computers will have a significant impact on today's security infrastructure. With the advent of powerful quantum computers public key cryptographic schemes will become vulnerable to Shor's quantum algorithm, undermining the security current communications systems. Post-quantum (or quantum-resistant) cryptography is an active research area, endeavoring to develop novel and quantum resistant public key cryptography. Amongst the various classes of quantum-resistant cryptography schemes, lattice-based cryptography is emerging as one of the most viable options. Its efficient implementation on software and on commodity hardware has already been shown to compete and even excel the performance of current classical security public-key schemes. This work discusses the next step in terms of their practical deployment, i.e., addressing the physical security of lattice-based cryptographic implementations. We survey the state-of-the-art in terms of side channel attacks (SCA), both invasive and passive attacks, and proposed countermeasures. Although the weaknesses exposed have led to countermeasures for these schemes, the cost, practicality and effectiveness of these on multiple implementation platforms, however, remains under-studied. Ayesha Khalid, Tobias Oder, Felipe Valencia, Máire O'Neill, Tim Güneysu, Francesco Regazzoni 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Compact, Scalable, and Efficient Discrete Gaussian Samplers for Lattice-Based CryptographyabstractLattice-based cryptography, one of the leading candidates for post-quantum security, relies heavily on discrete Gaussian samplers to provide necessary uncertainty, obfuscating computations on secret information. For reconfigurable hardware, the cumulative distribution table (CDT) scheme has previously been shown to achieve the highest throughput and the smallest resource utilisation, easily outperforming other existing samplers. However, the CDT sampler does not scale well. In fact, for large parameters, the lookup tables required are far too large to be practically implemented. This research proposes a hierarchy of multiple smaller samplers, extending the Gaussian convolution lemma to compute optimal parameters, where the individual samplers require much smaller lookup tables. A large range of parameter sets, covering encryption, signatures, and key exchange are evaluated. Hardware-optimised parameters are formulated and a practical implementation on Xilinx Artix-7 FPGA device is realised. The proposed sampling designs demonstrate promising performance on reconfigurable hardware, even for large parameters, that were otherwise thought infeasible. Ayesha Khalid, James Howe, Ciara Rafferty, Francesco Regazzoni 0001, Máire O'Neill |
ISCAS | 1 |
| 2018 | On Practical Discrete Gaussian Samplers for Lattice-Based CryptographyabstractLattice-based cryptography is one of the most promising branches of quantum resilient cryptography, offering versatility and efficiency. Discrete Gaussian samplers are a core building block in most, if not all, lattice-based cryptosystems, and optimised samplers are desirable both for high-speed and low-area applications. Due to the inherent structure of existing discrete Gaussian sampling methods, lattice-based cryptosystems are vulnerable to side-channel attacks, such as timing analysis. In this paper, the first comprehensive evaluation of discrete Gaussian samplers in hardware is presented, targeting FPGA devices. Novel optimised discrete Gaussian sampler hardware architectures are proposed for the main sampling techniques. An independent-time design of each of the samplers is presented, offering security against side-channel timing attacks, including the first proposed constant-time Bernoulli, Knuth-Yao, and discrete Ziggurat sampler hardware designs. For a balanced performance, the Cumulative Distribution Table (CDT) sampler is recommended, with the proposed hardware CDT design achieving a throughput of 59.4 million samples per second for encryption, utilising just 43 slices on a Virtex 6 FPGA and 16.3 million samples per second for signatures with 179 slices on a Spartan 6 device. James Howe, Ayesha Khalid, Ciara Rafferty, Francesco Regazzoni 0001, Máire O'Neill |
IEEE Trans. Computers | 2 |
| 2017 | Compact and provably secure lattice-based signatures in hardwareabstractLattice-based cryptography is a quantum-safe alternative to existing classical asymmetric cryptography, such as RSA and ECC, which may be vulnerable to future attacks in the event of the creation of a viable quantum computer. The efficiency of lattice-based cryptography has improved over recent years, but there has been relatively little investigation into hardware designs of digital signature schemes. In this paper, the first hardware design of the provably secure Ring-LWE digital signature scheme, Ring-TESLA, is presented, targeting a Xilinx Spartan-6 FPGA. The results better compactness of all previous lattice-based digital signature schemes in hardware, and can achieve between 104-785 signatures and 102-776 verifications per second. James Howe, Ciara Rafferty, Ayesha Khalid, Máire O'Neill |
ISCAS | 3 |
| 2017 | RC4-AccSuite: A Hardware Acceleration Suite for RC4-Like Stream CiphersabstractWe present RC4-AccSuite, a hardware accelerator, which combines the flexibility of an application specific instruction set processor and the performance of an application specific IC for the most widely deployed commercial stream cipher RC4 and its eight prominent variants, including Spritz (CRYPTO-2014 Rump-session). Our carefully designed instruction set architecture reuses combinational and sequential logic at its various pipeline stages and memories, saving up to 41% in terms of area, compared with the individual cores, while the power budget being dictated primarily by the variant used. Moreover, using state replication, noticeable throughput performance enhancement in RC4 variants is achieved. RC4-AccSuite possesses extensibility for future variants of RC4 with little or no tweaking. Ayesha Khalid, Goutam Paul 0001, Anupam Chattopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Time-independent discrete Gaussian sampling for post-quantum cryptographyabstractAs the development of a viable quantum computer nears, existing widely used public-key cryptosystems, such as RSA, will no longer be secure. Thus, significant effort is being invested into post-quantum cryptography (PQC). Lattice-based cryptography (LBC) is one such promising area of PQC, which offers versatile, efficient, and high performance security services. However, the vulnerabilities of these implementations against side-channel attacks (SCA) remain significantly understudied. Most, if not all, lattice-based cryptosystems require noise samples generated from a discrete Gaussian distribution, and a successful timing analysis attack can render the whole cryptosystem broken, making the discrete Gaussian sampler the most vulnerable module to SCA. This research proposes countermeasures against timing information leakage with FPGA-based designs of the CDT-based discrete Gaussian samplers with constant response time, targeting encryption and signature scheme parameters. The proposed designs are compared against the state-of-the-art and are shown to significantly outperform existing implementations. For encryption, the proposed sampler is 9× faster in comparison to the only other existing time-independent CDT sampler design. For signatures, the first time-independent CDT sampler in hardware is proposed. Ayesha Khalid, James Howe, Ciara Rafferty, Máire O'Neill |
FPT | 1 |
| 2016 | RunStream: A High-Level Rapid Prototyping Framework for Stream CiphersabstractWe present RunStream, a rapid prototyping framework for realizing stream cipher implementations based on algorithmic specifications and architectural customizations desired by the users. In the dynamic world of cryptography where newer recommendations are frequently proposed, the need of such tools is imperative. It carries out design validation and generates an optimized software implementation and a synthesizable Register Transfer Level Verilog description. Our framework enables speedy benchmarking against critical resources like area, throughput, power, and latency and allows exploration of alternatives. Using RunStream, we successfully implemented various stream ciphers and benchmarked the quality of results to be at par with published hand-optimized implementations. Ayesha Khalid, Goutam Paul 0001, Anupam Chattopadhyay, Faezeh Abediostad, Syed Imad Ud Din, Baishik Biswas, Prasanna Ravi |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2015 | New ASIC/FPGA Cost Estimates for SHA-1 CollisionsabstractSHA-1 remains, till date, the most widely used hash function, in spite of several successful cryptanalytic attacks against it. These attacks, however, remain impractical due to high computation complexity and associated cost. We endeavor to do cost-time product estimation for an attack by the aid of application-specific hardware acceleration. This work proposes an Application-Specific Instruction-set Processor (ASIP), named Cracken. Cracken is aimed to efficiently realize near collision attack on SHA-1. The estimations of the physical attack complexity is done using 65nm standard CMOS technology and commercial FPGA devices. It is estimated, with post-layout simulations, that Stevens' differential attack with an estimated complexity of 2^57.5, can be executed in 46 days using 4096 Cracken cores at a cost of Euros 15m. Estimation for real collision with complexity 2^61 is also done. Our cost-time estimates reveal that an FPGA-based attack is more efficient compared to ASIC. Previously reported SHA-1 attacks based on ASIC and cloud computing platforms are also compiled and benchmarked for reference. Ayesha Khalid, Anupam Chattopadhyay, Christian Rechberger, Tim Güneysu, Christof Paar |
DSD | 2 |
| 2013 | CoARX: a coprocessor for ARX-based cryptographic algorithmsabstractCryptographic coprocessors are inherent part of modern System-on-Chips. It serves dual purpose - efficient execution of cryptographic kernels and supporting protocols for preventing IP-piracy. Flexibility in such coprocessors is required to provide protection against emerging cryptanalytic schemes and to support different cryptographic functions like encryption and authentication. In this context, a novel crypto-coprocessor, named CoARX, supporting multiple cryptographic algorithms based on Addition (A), Rotation (R) and eXclusive-or (X) operations is proposed. CoARX supports diverse ARX-based cryptographic primitives. We show that compared to dedicated hardware implementations and general-purpose microprocessors, it offers excellent performance-flexibility trade-off including adaptability to resist generic cryptanalysis. Khawar Shahzad, Ayesha Khalid, Zoltán Endre Rákossy, Goutam Paul 0001, Anupam Chattopadhyay |
DAC | 2 |
| 2013 | SI-DFA: Sub-expression integrated Deterministic Finite Automata for Deep Packet InspectionabstractFinite automata is widely used for Deep Packet Inspection (DPI) of network traffic. Two types of automata employed for this purpose are Non-deterministic Finite Automata (NFA) and Deterministic Finite Automata (DFA). An NFA suffers from a large memory bandwidth per character due to multiple active states. A DFA, in comparison, ensures a linear processing time of O(1) for memory based architectures. However, the DFA state explosion conditions commonly occurring in today's NIDS rule-sets, render the automata with practically infeasible memory space requirements. To avoid state blowup we propose a semi-deterministic automata, Sub-expression Integrated DFA (SI-DFA), that ensures processing time of a single standard DFA. Rules are broken into sub-expressions at blowup conditions and compiled into a single DFA along with an association table, to correctly encapsulate equivalent automata. We list the rare cases in regular expressions for which sub-expression Integration is incorrect and present methodology to detect their occurrences. We evaluate SI-DFA on real-world rule-sets like Bro, Snort and Linux filters and compare their performance with the state-of-the-art hybrid automata solutions. SI-DFA renders a 66-97% reduction in processing bandwidth, up to 68% lower space requirement and an improvement trend with increasing rule complexity when compared to the traditional solutions. Ayesha Khalid, Rajat Sen, Anupam Chattopadhyay |
HPSR | 1 |
| 2012 | Designing high-throughput hardware accelerator for stream cipher HC-128abstractDue to ubiquitous deployment of embedded systems, security and privacy are emerging as major design concerns and new stream ciphers are being proposed by the cryptographic community. HC-128 is one of the recent stream ciphers that received attention after its selection as an eStream candidate. Till date, the cipher is believed to have a good security margin. In this paper we study several implementation issues for HC-128 in a disciplined manner. We first discuss the experience on embedded and customizable processors. Then we consider a dedicated hardware accelerator implementation. Further we explore several parallelization strategies for improving throughput. To the best of our knowledge such a detailed implementation exercise has not been presented in the literature. Our novel implementation strategies mark the fastest HC-128 execution reported till date. Anupam Chattopadhyay, Ayesha Khalid, Subhamoy Maitra, Shashwat Raizada |
ISCAS | 2 |