VLDB 2026 Research / reviewers in the wild / expert
Ziying Ni
dblp:275/4805
· DBLP profile ↗
15ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0001-8300-8865ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 9 first-author · 14 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimized NTT Architecture Based on the Plantard Algorithm for ML-KEM and ML-DSAabstractModular multiplication is a vital operation in the Number Theoretic Transform (NTT), significantly enhancing polynomial multiplication in Post-Quantum Cryptography (PQC). The design efficiency of modular multiplication directly influences the computational performance of polynomial computation units. This work marks the first hardware-oriented improvement of the Plantard algorithm, optimizing the NTT architecture. We modify the Plantard algorithm and propose three innovative enhanced versions tailored for lattice-based cryptography (LBC). By employing pre-processed twiddle factors for result correction and eliminating an additional constant multiplication, we greatly simplify the computation steps. Based on these improvements, we further design a lightweight BRAM-free iterative NTT and a high-speed Multi-path Delay Commutator (MDC) pipelined NTT, both targeting the ML-KEM and ML-DSA parameter sets. Implementation results on the Xilinx Artix-7 platform demonstrate that our Plantard_preω design reduces slice usage by 22% to 53% and delay by 41.1% to 44.3% compared to existing algorithms like Barrett and K2RED. Furthermore, our iterative NTT design achieves the minimal area-time product (ATP) among state-of-the-art implementations, with reductions of 61.3% and 40.8% in ENS, and frequency increases of 72.7% and 145.4% for ML-KEM and ML-DSA, respectively. The pipelined NTT design also reduces delay by 23.7% and ATP by 4.9%, showcasing the compactness and superior performance of our approach. Bei Wang 0013, Ziying Ni, Mengxue Li, Fei Lyu 0002, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 3 |
| 2026 | LightHD: A Lightweight and High-Performance Hardware Accelerator of CRYSTALS-DilithiumabstractCRYSTALS-Dilithium serves as the foundation of the NIST-standardised PQC digital signature scheme, and has been declared as the first recommended digital signature algorithm. However, due to the computational complexity and intricate processing flow of CRYSTALS-Dilithium, two limitations are shown in existing methods: its applicability on resource-constrained devices is limited and the performance reported so far remains relatively low. This paper presents a lightweight yet high-performance hardware architecture that optimises the core computational units of CRYSTALS-Dilithium. First, an iterative dual-Keccak SHA-3 module is proposed, where two cores operate with a 26-cycle offset to accelerate processing without compromising frequency. In addition, the rejection sampler is streamlined by two compact registers for intermediate values and counters, improving efficiency when consuming interleaved SHA-3 outputs. Second, for small bit-width polynomials, we eliminate the first NTT stage via lookup tables and data regrouping, reducing NTT cycles by 11.7% with little hardware overhead. Further hardware savings are achieved by maximising IP core utilisation and simplifying input multiplexers. Furthermore, a compact scheduling strategy ensures that all intermediate storage fits within a single polynomial-sized memory block. On Xilinx Artix-7 FPGAs, the design reduces hardware overhead by 14.2% compared with state-of-the-art lightweight implementations. Across three security levels, KeyGen and Verify are 27.3% and 13.5% faster, respectively, than high-performance prior designs. At level 5, the best-case Sign latency is only 120 µs. Ziying Ni, Ayesha Khalid, Zhaoyu Zhang 0001, Yijun Cui, Weiqiang Liu 0001, Máire O'Neill |
IEEE Trans. Computers | 1 |
| 2025 | AxRA: Approximate Rowhammer Attack for Modern DRAM SystemsabstractApproximate computing achieves high performance or less power consumption in various fault-tolerant applications, e.g., image processing, artificial intelligence (AI), etc. However, the introduction of approximate computing brings new security vulnerabilities, which threaten the entire computing system. In this paper, a novel Rowhammer attack is proposed, which utilises the approximate data stored in DRAM memories to achieve higher attack effectiveness. Compared to Rowhammer attack to DRAM memory without approximate data, the proposed method achieves more bit-flips resulting in significant data corruption. The proposed attack is implemented and evaluated on DRAM chips with a real user case, object detection using neural network. The accuracy of detection on the baseline image is employed to verify the impact of proposed attack approach. The results show that the proposed Rowhammer attack with approximate data introduces extra 33% bit-flips on victim rows than a conventional Rowhammer attack without approximate data. It also introduces up to ∼75% accuracy reduction of MNIST neural network proportionally to the increment of attack activation number. Yuhang Hao, Yun Wu 0003, Ziying Ni, Jack Miskelly, Máire O'Neill, Chongyan Gu |
ISCAS | 3 |
| 2025 | Ultra-compact and Side-channel Resistant Design of FIFO-based NTT Core for PQCsabstractCryptographic algorithms like CRYSTALS-Kyber and Dilithium might be insecure with their naive implementation facing side-channel attacks (SCA). This work presents a compact implementation of Number Theoretic Transform (NTT) with shuffling countermeasure against power analysis attacks (PA). At first, a compact FIFO-only Shuffler module is presented to perform group-wise first-index randomization (FIR). A modified butterfly (BF) unit using optimized modulus reduction is then promoted to restore the misaligned data flow, which is critical for forming efficient shuffle pattern. The shuffler module and BF unit are then used in a pipelined BRAM-free NTT baseline. Through efficient shuffling, the proposed design maintains compactness akin to its baseline while enhancing robust SCA resistance with a permutation space of up to 2494. Compared to its prior state-of-the-art designs, the proposed NTT core presents an improvement of 53.2% in area-time trade-off while offering ×1.7 times more bits of randomness to improve hardware security. Jiatong Tian, Yijun Cui, Ziying Ni, Bei Wang 0013, Fei Lyv, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2025 | Instruction-Based High-Performance Hardware Controller of CRYSTALS-Kyber With Balanced Resource UtilizationabstractPost-quantum cryptography (PQC) aims to ensure information security in the era following the emergence of quantum computers. Lattice-based cryptography (LBC) algorithms have shown significant promise in the standardization process of post-quantum cryptography. This paper proposes an instruction-based high-performance hardware controller of CRYSTALS-Kyber. By designing a highly flexible instruction-based architecture, the control unit evenly distributes instructions and enables independent control of internal modules, significantly enhancing the scalability and adaptability of the hardware. Additionally, the integration of a reconfigurable polynomial operation array (RPOA) unit and optimization of data storage formats further improve computational efficiency and resource utilization. Implementation results on Artix-7 FPGA show that the architecture operates at a frequency exceeding 300 MHz, achieving a performance improvement of 41.3% to 170% compared to the latest designs, while significantly reducing resource overhead. The resource costs for the three security levels are 8112 LUTs, 6077 FFs, and 2523 SLICEs, respectively, with overall computation times of$34.7~\mu s$,$53.4~\mu s$, and$78.5~\mu s$. The proposed design demonstrates outstanding performance, resource efficiency, and energy consumption, providing an efficient and cost-effective hardware solution for the practical deployment of post-quantum cryptography. Yijun Cui, Ziying Ni, Zhuoyao Zhang, Chenghua Wang, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | A Highly Hardware Efficient ML-KEM Accelerator with Optimised Architectural LayersabstractThe Module-Lattice-Based Key encapsulation Mechanism (ML-KEM) scheme, which is currently being standardised, is a quantum attack resistant KEM that is based on CRYSTALS-Kyber. CRYSTALS-Kyber is the only Public-key Encryption (PKE)/ KEM scheme selected in the first set of successful candidates as part of the NIST initiated Post-Quantum Cryptography (PQC) process. ML-KEM scheme includes three different security levels, namely security level 1, 3, and 5. In this research, we propose a highly area-time efficient hardware ML-KEM architecture. The architecture comprises three computational layers. The first layer comprises a hash and sampling module; the second layer includes a number theoretic transform (NTT), its inverse (INTT) and a point-wise multiplication (PWM) module; and the third layer comprises addition, compressing and encoding. Intra-layer pipelining and out-of-layer scheduling ensures that either layer 1 or layer 2 operate in the shortest time. In the reduction module, we propose a novel hybrid architecture to obtain the final result within 2 cycles with low area consumption. In the NTT module, the PWM pipelining method is modified and an optimised iterative FIFO access method is adopted to reduce the size of FIFO units by 55% over previous research. Look-up tables are also used to replace the first-stage of the NTT to reduce 8 cycles. Furthermore, the memory unit uses only FIFOs, the size are optimised based on the requirements of the most resource-intensive function in ML-KEM (ML-KEM.CPA.Dec). The results show that the proposed architecture has a 48.2%, 41.2%, and 78.1% reduction in computational time in comparison to previous work for security levels 1, 3, and 5, respectively. In addition, the area of proposed optimised ML-KEM designs is reduced by 73%, 70%, 76% and resulting in an improved area-time (AT) product of 15.8%, 10.7%, and 11.3%, for the Level 1, 3, and 5 security levels respectively, compared with state-of-the-art designs. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | Bitstream Fault Injection Attacks on CRYSTALS Kyber Implementations on FPGAsabstractCRYSTALS-Kyber is the only Public-key Encryption (PKE)/ Key-encapsulation Mechanism (KEM) scheme that was chosen for standardization by the National Institute of Standards and Technology initiated Post-quantum Cryptography competition (so called NIST PQC). In this paper, we show the first successfully malicious modifications of the bitstream of a Kyber FPGA implementation. We successfully demonstrate 4 different attacks on Kyber hardware implementations on Artix-7 FPGAs that either reduce the complexity of polynomial multiplication operations or enable direct secret key/ message recovery by: disabling BRAMs, disabling DSPs, zeroing NTT ROM and tampering with CBD2 results. Two of our attacks are generic in nature and the other two require reverse-engineering or a detailed knowledge of the design. We evaluate the feasibility of the four attacks, among which the zeroing NTT ROM and tampering with the CBD2 result attacks produce higher public key and ciphertext complexity and thus are difficult to be detected. Two countermeasures are proposed to prevent the attacks proposed in this paper. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
DATE | 1 |
| 2024 | FPGA Bitstream Fault Injection Attack and Countermeasures on the Sampling Counter in CRYSTALS KyberabstractThe CRYSTALS Kyber algorithm is the public key encryption (PKE)/ key encapsulation mechanism (KEM) protocol undertaken for standardization by the US National Institute of Standards and Technology (NIST) the PQC competition and serves as the foundation for the Module-Lattice-Based (ML)-KEM scheme. The inherently strong security properties of the Kyber algorithm are considered to be resistant to attacks under quantum computers, but the security of its FPGA-based hardware implementation circuitry is still worth considering. In this work, we introduce the Nonce counter disabling attack, which targets the binomial distribution sampling process. We demonstrate that, in the modified primes of Kyber from Round 2, it is also effectively deduce the secret key s by equating it with the noise e. Our implementation of this attack on a Nexys 4 FPGA, with an additional DSP disabling filtering process to pinpoint the LUT. This attack is applicable to both the key generation and key encapsulation phases, and only need to modify 32-bit bitstream. Finally, We propose the Nonce counter check and the splitting of the Nonce computation cycles methods to to prevent this attack in hardware design-level. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 1 |
| 2023 | Towards a Lightweight CRYSTALS-Kyber in FPGAs: an Ultra-lightweight BRAM-free NTT CoreabstractCRYSTALS-Kyber is the first quantum-resilient, lattice-based Public Key Encryption (PKE)/Key Encapsulation Mechanism (KEM) cryptosystem that is chosen by the ongoing National Institute of Standards and Technology post-quantum cryptography standardization (NIST PQC) for standardization. This work presents a lightweight and efficient, FPGA-based hardware implementation for polynomial multiplication unit (NTT), which is the major bottleneck in the Kyber scheme. As a first step, an optimzed modular multiplication architecture combining KRED and lookup table-based algorithms is presented, which reduces the resources of slices by 16.7%. It is used in a pipelined NTT/INTT architecture that is completely BRAM free and instead uses 3 FIFOs for coefficients storage. We hereby present the most compact FPGA based design for NTT architecture in Kyber till date. Experimental results bench marked on comparable FPGA devices show that our proposed design is 36-75% better than the state-of-the-art implementations in terms of hardware efficiency for NTT/INTT calculations and$3.4-4.4\times$better for the Point-wise Multiplication (PWM) operation. Ziying Ni, Ayesha Khalid, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 1 |
| 2023 | HPKA: A High-Performance CRYSTALS-Kyber Accelerator Exploring Efficient PipeliningabstractCRYSTALS-Kyber (Kyber) was recently chosen as the first quantum resistant Key Encapsulation Mechanism (KEM) scheme for standardisation, after three rounds of the National Institute of Standards and Technology (NIST) initiated PQC competition which begin in 2016 and search of the best quantum resistant KEMs and digital signatures. Kyber is based on the Module-Learning with Errors (M-LWE) class of Lattice-based Cryptography, that is known to manifest efficiently on FPGAs. This work explores several architectural optimizations and proposes a high-performance and area-time (AT) product efficient hardware accelerator for Kyber. The proposed architectural optimizations include inter-module and intra-module pipelining, that are designed and balanced via FIFO based buffering to ensure maximum parallelisation. The implementation results show that compared to state-of-the-art designs, the proposed architecture delivers 25–51% speedups for Kyber's three different security levels on Artix-7 and Zynq UltraScale+ devices, and a 50–75% reduction in DSPs at comparable security level. Consequently, the proposed design achieve higher AT product efficiencies of 19–33%. Ziying Ni, Ayesha Khalid, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Computers | 1 |
| 2022 | High Performance FPGA-based Post Quantum Cryptography ImplementationsabstractPost-quantum Cryptography (PQC) is an umbrella term for cryptographic schemes based on hard mathematical problems which are resistant to attacks by quantum computers. The National Institute of Standards and Technology (NIST) initiated a PQC standardisation process in 2017, with a total of 4 algorithms selected for standardisation after round 3 and 4 undertaken for further analysis in Round 4 in 2022. PQC schemes on hardware devices, such as Field Programmable Gate Arrays (FPGA), show the potential of higher throughput performance, for comparable security, at the cost of high area and power consumption. The major aim of this thesis is to help facilitate the global transition to a post quantum secure set of security protocols. This thesis will focus on the optimisation of the the hardware architectures to improve the computational speed and reduce the area overhead. The side channel analysis vulnerabilities and their countermeasures will also be studied. Ziying Ni, Ayesha Khalid, Máire O'Neill |
FPL | 1 |
| 2022 | A Lightweight and Efficient Schoolbook Polynomial Multiplier for SaberabstractSaber is a lattice-based post-quantum cryptography (PQC) algorithm, which is still a candidate in the 3rdRound of National Institute of Standards and Technology (NIST) PQC standardization process. Saber provides a great advantage of being lightest among all the candidates, so a suitable choice for resource-constraint platforms. Polynomial multiplication occupies most of the resources in hardware implementation of Saber, which needs to be optimized for the efficient hardware implementation. In this work, a lightweight and efficient schoolbook polynomial multiplier is proposed. The architecture includes an efficient multiplication strategy that compute four coefficient-wise multiplication per cycle along with the multiplication operand loading technique being designed for the compact multiplier. The proposed multiplier on Artix-7 FPGA, achieves a frequency of 130 MHz and fits into 201 slices. Compared with the state-of-the-art lightweight schoolbook implementations for Saber, our design has a 30% improved frequency and saves 15.8% of the clock counts at the cost of only 3.7% more LUTs. Yuantuo Zhang, Yijun Cui, Ziying Ni, Dur-e-Shahwar Kundi, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2022 | A High-Performance SIKE Hardware AcceleratorabstractSupersingular isogeny key encapsulation (SIKE) is a promising candidate in the NIST postquantum cryptography (PQC) standardization process, which has the smallest key lengths. It is the only isogeny-based cryptographic scheme in the NIST list that leverages the traditional elliptic curve cryptography (ECC) arithmetic; however, the high computational complexity is one of its limiting factors. In this work, we proposed a high-performance hardware architecture for the SIKE protocol. The architecture includes an improved multiplier based on the high-performance finite field multiplication (HFFM) algorithm which is 15%–20.7% faster than the previous multiplier based on the HFFM algorithm and a unified adder/subtractor with radix$3^{b}$. In addition, it also comprises an efficient scheduler strategy that decomposes all the functions of SIKE into finite field$F_{p}$and then effectively schedules through optimized multiplication chains for maximal performance. The proposed architecture is synthesized and implemented on Xilinx Virtex-7 FPGA for all the four variants of SIKE having security levels from 1 to 5 and achieved 2.6%–7.8% faster speeds as well as consumed less equivalent number of slices (ENS) than the state-of-the-art designs. In the comparison of area and time (AT), the proposed architecture is 14.2%–34.5% lower than the previous architecture. Ziying Ni, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | High-Performance Systolic Array Montgomery Multiplier for SIKEabstractIn theory, the speed of quantum computers is much faster than classical computers, which poses a threat to the Public Key Cryptography (PKC) that are currently in use. Post Quantum Cryptography (PQC) is a class of cryptography based on complex mathematical problems that are difficult to be attacked by quantum computers. The Supersingular Isogeny Key Encapsulation (SIKE) protocol is one of candidate algorithms for the US National Institute of Standards and Technology (NIST) PQC standardization process and survived to the Round 3. In this paper, we reconstruct the systolic array based Montgomery multiplier architecture for SIKE, using a three-stage pipeline that results in frequency improvement of 21.4%. The proposed multiplier consumed fewer DSP resources than the state-of-the-art SIKE designs and has a speed increase up to 12.7%. Ziying Ni, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 1 |
| 2020 | High Performance Modular Multiplication for SIDHabstractThe latest research indicates that quantum computers will be realized in the near future. In theory, the computation speed of a quantum computer is much faster than current computers, which will pose a serious threat to current cryptosystems. Post-quantum cryptography (PQC) is a class of cryptography based on underlying mathematical problems that are considered infeasible to crack even with access to a quantum computer. The supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol is a new post-quantum cryptosystem, which offers advantages in reduced secret key length and attack resistance. SIDH is the basis of the supersingular isogeny key encapsulation (SIKE) protocol, which is in the second round of the U.S. National Institute of Standards and Technology (NIST) PQC standardization process. In this article, we propose a new modular multiplication algorithm and a new interleaved hardware architecture for SIDH. Performance results for the proposed modular multiplier using four parameter sets for the prime, p that correspond to the SIKE Round 2 parameter sets show significant advantages in speed. Weiqiang Liu 0001, Ziying Ni, Jian Ni, Ciara Rafferty, Máire O'Neill |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |