Trong-Hung Nguyen

dblp:364/3618 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0001-0952-0534ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Compact FALCON FFT/NTT Accelerator for Post-Quantum Cryptography
abstract
FALCON is one of four algorithms selected by NIST to standardize post-quantum cryptography standards. FALCON is a digital signature algorithm based on NTRU lattice with difficulty based on the short vector problem. While Kyber and Dilithium algorithms are only based on NTT operations, FALCON uses both NTT and FFT, which is a barrier to Falcon’s hardware implementation. This paper proposes a compact architecture that supports FFT and NTT for the FALCON algorithm. First, we propose an architecture that executes floating-point and complex number operations with theoretic speed and low area requirements. Then, we design a processing element that performs FFT with complex number operations. Finally, we propose an NTT architecture that reuses the resources used for FFT execution with high parallelism. The FPGA implementation results show that the FFT execution takes 2.048k CCs and 4.608k CCs for the 512-point and 1024-point FFT/IFFT, respectively. The NTT/INTT operation takes 288 CCs for FALCON-512 and 640 CCs for FALCON-1024. The speedup improves from 3× to 9.6× for FFT and up to 18× for NTT implementations compared to previous studies.
Duc-Thuan Dam, Thai-Ha Tran, Trong-Hung Nguyen, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS3
2025 A Low-Latency Polynomial Arithmetic Unit for ML-KEM and ML-DSA Standards
abstract
Existing communication protocols based on public key cryptography (PKC) functions will no longer be secure in the quantum era. NIST has released standards for key encapsulation and digital signature mechanisms based on module lattice (ML-KEM and ML-DSA) to address this challenge. In this paper, we propose a unique, high-performance arithmetic unit capable of performing all the polynomial operations needed for ML-KEM and ML-DSA (KDA). The proposed KDA architecture includes a computational unit that supports one 4×1 NTT configuration for ML-DSA and double 4×1 NTT configurations for ML-KEM. A two-step NTT data flow and configurable memory unit are introduced to reorder and store coefficients for all operations. Moreover, we propose the re-used twiddle factor method for NTT and point-wise multiplication. We have implemented three design versions, ML-KEM standalone, ML-DSA standalone, and KDA, and compared them to the existing studies. The comparison shows that our KDA achieves superior ATP performance, improving ×1.1-×3.6.
Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS1
2025 A Timing-Constrained Design Methodology for Radix- 2k NTT in Polynomial Arithmetic
abstract
Polynomial modular multiplication is the most complex and costly operation in homomorphic encryption (HE) and post-quantum cryptography (PQC). Using the Number Theoretic Transform (NTT) helps reduce the complexity of multiplication to quasi-linear O($N\,\textup{log}_{2}N$). Although NTT significantly impacts the performance of HE and PQC, existing NTT-based multipliers often fall short due to inefficient data movement and large memory overhead. Notably, deploying low-latency cryptosystems incurs more significant costs with reduced acceleration gains. To overcome these constraints, we introduce a pioneering methodology called timing-constrained NTT (TCO-NTT). We propose an innovative time-controlled memory (TCM) structure that re-orders and stores coefficients within each stage of the NTT. Then, we employ the divide-and-conquer strategy, allowing freely configurable parallelism levels. Besides, our proposed methodology can generalize to radix-2kNTT and supports any arbitrary polynomial degreeNand scale factorpvalues. We evaluate the proposed TCO-NTT on typical HE and PQC parameter sets across multiple levels of parallelism and radix-2kNTT configurations. FPGA implementation results demonstrate that our TCO-NTT achieves minimal hardware cost while consistently executing the NTT in a near-theoretical execution time. Our area-time product (ATP) reports about LUT-ATP (LATP), FF-ATP (FATP), and BRAM-ATP (BATP) surpass the reported-to-date NTT designs by up to 10.2×, 17.8× and 47.2×. The proposed TCO-NTT sets new records for NTT-based multiplier efficiency, laying the foundation for implementing HE and PQC in real-time applications.
Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Tuan-Kiet Dang, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 Efficient Hardware Implementation of the Lightweight CRYSTALS-Kyber
abstract
Quantum computing raises questions about the security of data encrypted using modern methods. Hence, the National Institute of Standards and Technology (NIST) has undertaken standardization of post-quantum cryptography (PQC) algorithms to defend against attacks from both classical and quantum computers. Following four rounds of evaluation, CRYSTALS-Kyber has been selected for standardization. In this paper, we present an efficient hardware architecture of CRYSTALS-Kyber for resource-constrained IoT devices. Firstly, we propose a compact hash module for CRYSTALS-Kyber. A single buffer is designed to perform padding, hashing, and holding data. Hence, using large FIFOs for data input/output is eliminated. Then, we propose a novel non-memory-based iterative number theoretic transform (NMI-NTT) architecture. Finally, the data flow between modules is optimized to improve parallelization and execution time. Implementation results on an Artix-7 FPGA show that our design consumes minimal hardware resources compared to the designs reported to date, corresponding to 5487 LUTs, 3426 FFs, 1548 SLICEs, 3.5 BRAMs, and 2 DSPs. Our design computes key generation, encapsulation, and decapsulation phases in 3.3/4.5/6.1 K-cycles for Kyber512, 5.6/7.1/9.2 K-cycles for Kyber768, and 8.5/10.1/12.9 K-cycles for Kyber1024, with 185MHz operating frequency. Our area-time-product (ATP) performance outperforms other designs.
Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Binh Kieu-Do-Nguyen, Cong-Kha Pham, Trong-Thuc Hoang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 Unified-pipelined NTT Architecture for Polynomial Multiplication in Lattice-based Cryptosystems
abstract
Number Theoretic Transformation (NTT) is commonly employed to speed up polynomial multiplication in post-quantum Lattice-Based Cryptography (LBC). A current trend in NTT hardware design involves using an iterative approach for forward and inverse NTT (INTT) computations. However, this iterative method demands substantial temporary memory and complex memory access patterns. This paper introduces a unified-pipelined NTT architecture for high-performance LBC cryptosystems. Our butterfly units employ a specially crafted Digital Signal Processing (DSP) for modular integer multiplication. Consequently, NTT and INTT calculations are carried out more swiftly with minimal hardware requirements, eliminating the need for DSP and Block Random Access Memory (BRAM). We applied this novel architecture to various parameter sets of LBC and implemented it on the Xilinx FPGA platform for comparison with state-of-the-art studies. Implementation results show that the proposed NTT architectures have outstanding hardware area and operating frequency improvements. The Area Time Product (ATP) is significantly improved, equivalent to at least 53% to 94% compared to the best designs reported to date.
Trong-Hung Nguyen, Nguyen The Binh, Huynh Phuc Nghi, Cong-Kha Pham, Trong-Thuc Hoang
ISCAS1
2024 An Efficient Method for Accelerating Kyber and Dilithium Post-Quantum Cryptography
abstract
Post-quantum cryptography (PQC) algorithms were introduced in response to the threats of attacks using quantum computers. The CRYSTALS-Kyber and CRYSTALS-Dilithium are two of the algorithms chosen by NIST to standardize the PQC, which are lattice-based algorithms. Number theoretic transform (NTT) helps lattice-based algorithms reduce latency, but it is still their bottleneck. Along with that, the RISC-V instruction set architecture also opens up flexible methods to solve different problems. This paper proposes a RISC-V system-on-a-chip (SoC) architecture with a computational accelerator for NTT-based calculations for Kyber and Dilithium. Implementation results show that software running on proposed SoC using accelerators has improved in NTT/INTT by up to$36.75\times/42.69\times$compared to software on embedded devices, up to$4.07\times/4.38\times$for software running on RISC-V SoCs, and up to$8.11\times$for NTT of the previous software/hardware architectures.
Duc-Thuan Dam, Trong-Hung Nguyen, Thai-Ha Tran, Binh Kieu-Do-Nguyen, Trong-Thuc Hoang, Cong-Kha Pham
PST2