EDBT 2026 Demo / reviewers in the wild / expert
Duc-Thuan Dam
dblp:364/3419
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-8998-9721ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Countering Side-Channel Attacks With a Dynamic S-Box Based on Affine Transformations and Gold SequencesabstractAdvanced cryptographic devices employ multiscale countermeasures to bolster resilience against side-channel analysis (SCA). In masking-based defenses, secure substitution-boxes (S-boxes) and effective masking schemes are paramount. Additionally, the time-based hiding techniques, leveraging multiple clocks for individual encryption operations, offer significant protection. This article introduces a novel multiscale countermeasure: an improved tower field masking scheme integrated with an affine transformation-based dynamic S-box. Crucially, we incorporate Gold sequences to generate both a random clock source for horizontal hiding and random values for masking. Extensive evaluation using up to five million power traces demonstrates the robustness of our approach against standard correlation power analysis (CPA) and alignment preprocessing techniques, including sliding window and amplitude peak localization. Experimental results show a measurement-to-disclosure (MTD) improvement of at least$150\times $compared to unprotected implementations using stand-alone masking and$375\times $with our multiscale approach. Furthermore, we demonstrate resilience against recent robust profiled deep learning SCA, which could only recover four subkeys even with one million traces. Thai-Ha Tran, Duc-Thuan Dam, Tuan-Kiet Dang, Duc-Hung Le, Trong-Thuc Hoang, Cong-Kha Pham |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Compact FALCON FFT/NTT Accelerator for Post-Quantum CryptographyabstractFALCON is one of four algorithms selected by NIST to standardize post-quantum cryptography standards. FALCON is a digital signature algorithm based on NTRU lattice with difficulty based on the short vector problem. While Kyber and Dilithium algorithms are only based on NTT operations, FALCON uses both NTT and FFT, which is a barrier to Falcon’s hardware implementation. This paper proposes a compact architecture that supports FFT and NTT for the FALCON algorithm. First, we propose an architecture that executes floating-point and complex number operations with theoretic speed and low area requirements. Then, we design a processing element that performs FFT with complex number operations. Finally, we propose an NTT architecture that reuses the resources used for FFT execution with high parallelism. The FPGA implementation results show that the FFT execution takes 2.048k CCs and 4.608k CCs for the 512-point and 1024-point FFT/IFFT, respectively. The NTT/INTT operation takes 288 CCs for FALCON-512 and 640 CCs for FALCON-1024. The speedup improves from 3× to 9.6× for FFT and up to 18× for NTT implementations compared to previous studies. Duc-Thuan Dam, Thai-Ha Tran, Trong-Hung Nguyen, Trong-Thuc Hoang, Cong-Kha Pham |
ISCAS | 1 |
| 2025 | A Low-Latency Polynomial Arithmetic Unit for ML-KEM and ML-DSA StandardsabstractExisting communication protocols based on public key cryptography (PKC) functions will no longer be secure in the quantum era. NIST has released standards for key encapsulation and digital signature mechanisms based on module lattice (ML-KEM and ML-DSA) to address this challenge. In this paper, we propose a unique, high-performance arithmetic unit capable of performing all the polynomial operations needed for ML-KEM and ML-DSA (KDA). The proposed KDA architecture includes a computational unit that supports one 4×1 NTT configuration for ML-DSA and double 4×1 NTT configurations for ML-KEM. A two-step NTT data flow and configurable memory unit are introduced to reorder and store coefficients for all operations. Moreover, we propose the re-used twiddle factor method for NTT and point-wise multiplication. We have implemented three design versions, ML-KEM standalone, ML-DSA standalone, and KDA, and compared them to the existing studies. The comparison shows that our KDA achieves superior ATP performance, improving ×1.1-×3.6. Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Trong-Thuc Hoang, Cong-Kha Pham |
ISCAS | 2 |
| 2025 | Enhanced Tower Field Mask Scheme with Affine Transformation-based Dynamic S-boxabstractMasking countermeasures are robust solutions applied to cryptographic devices to improve their side-channel analysis resistance. Implementing substitution boxes (S-boxes) and using an efficient mask scheme are essential for improving security in modern ciphers. Consequently, this paper proposes an improved tower field mask scheme with an affine transformation-based dynamic S-box. The approach resists both Correlation Power Analysis attacks with Hamming Weight and Hamming Distance models, even when employing up to two million power traces. The measurement-to-disclosure improvement for the extracted key byte is at least 158× higher than the previous scheme, while our hardware overhead is around 1.04×. Furthermore, the proposal enhances the devices’ resistance to recent Deep-Learning Side-Channel Analyses. Thai-Ha Tran, Duc-Thuan Dam, Van-Phuc Hoang, Trong-Thuc Hoang, Cong-Kha Pham |
ISCAS | 2 |
| 2025 | A Timing-Constrained Design Methodology for Radix- 2k NTT in Polynomial ArithmeticabstractPolynomial modular multiplication is the most complex and costly operation in homomorphic encryption (HE) and post-quantum cryptography (PQC). Using the Number Theoretic Transform (NTT) helps reduce the complexity of multiplication to quasi-linear O($N\,\textup{log}_{2}N$). Although NTT significantly impacts the performance of HE and PQC, existing NTT-based multipliers often fall short due to inefficient data movement and large memory overhead. Notably, deploying low-latency cryptosystems incurs more significant costs with reduced acceleration gains. To overcome these constraints, we introduce a pioneering methodology called timing-constrained NTT (TCO-NTT). We propose an innovative time-controlled memory (TCM) structure that re-orders and stores coefficients within each stage of the NTT. Then, we employ the divide-and-conquer strategy, allowing freely configurable parallelism levels. Besides, our proposed methodology can generalize to radix-2kNTT and supports any arbitrary polynomial degreeNand scale factorpvalues. We evaluate the proposed TCO-NTT on typical HE and PQC parameter sets across multiple levels of parallelism and radix-2kNTT configurations. FPGA implementation results demonstrate that our TCO-NTT achieves minimal hardware cost while consistently executing the NTT in a near-theoretical execution time. Our area-time product (ATP) reports about LUT-ATP (LATP), FF-ATP (FATP), and BRAM-ATP (BATP) surpass the reported-to-date NTT designs by up to 10.2×, 17.8× and 47.2×. The proposed TCO-NTT sets new records for NTT-based multiplier efficiency, laying the foundation for implementing HE and PQC in real-time applications. Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Tuan-Kiet Dang, Trong-Thuc Hoang, Cong-Kha Pham |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Efficient Hardware Implementation of the Lightweight CRYSTALS-KyberabstractQuantum computing raises questions about the security of data encrypted using modern methods. Hence, the National Institute of Standards and Technology (NIST) has undertaken standardization of post-quantum cryptography (PQC) algorithms to defend against attacks from both classical and quantum computers. Following four rounds of evaluation, CRYSTALS-Kyber has been selected for standardization. In this paper, we present an efficient hardware architecture of CRYSTALS-Kyber for resource-constrained IoT devices. Firstly, we propose a compact hash module for CRYSTALS-Kyber. A single buffer is designed to perform padding, hashing, and holding data. Hence, using large FIFOs for data input/output is eliminated. Then, we propose a novel non-memory-based iterative number theoretic transform (NMI-NTT) architecture. Finally, the data flow between modules is optimized to improve parallelization and execution time. Implementation results on an Artix-7 FPGA show that our design consumes minimal hardware resources compared to the designs reported to date, corresponding to 5487 LUTs, 3426 FFs, 1548 SLICEs, 3.5 BRAMs, and 2 DSPs. Our design computes key generation, encapsulation, and decapsulation phases in 3.3/4.5/6.1 K-cycles for Kyber512, 5.6/7.1/9.2 K-cycles for Kyber768, and 8.5/10.1/12.9 K-cycles for Kyber1024, with 185MHz operating frequency. Our area-time-product (ATP) performance outperforms other designs. Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Binh Kieu-Do-Nguyen, Cong-Kha Pham, Trong-Thuc Hoang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | An Efficient Method for Accelerating Kyber and Dilithium Post-Quantum CryptographyabstractPost-quantum cryptography (PQC) algorithms were introduced in response to the threats of attacks using quantum computers. The CRYSTALS-Kyber and CRYSTALS-Dilithium are two of the algorithms chosen by NIST to standardize the PQC, which are lattice-based algorithms. Number theoretic transform (NTT) helps lattice-based algorithms reduce latency, but it is still their bottleneck. Along with that, the RISC-V instruction set architecture also opens up flexible methods to solve different problems. This paper proposes a RISC-V system-on-a-chip (SoC) architecture with a computational accelerator for NTT-based calculations for Kyber and Dilithium. Implementation results show that software running on proposed SoC using accelerators has improved in NTT/INTT by up to$36.75\times/42.69\times$compared to software on embedded devices, up to$4.07\times/4.38\times$for software running on RISC-V SoCs, and up to$8.11\times$for NTT of the previous software/hardware architectures. Duc-Thuan Dam, Trong-Hung Nguyen, Thai-Ha Tran, Binh Kieu-Do-Nguyen, Trong-Thuc Hoang, Cong-Kha Pham |
PST | 1 |
| 2024 | Hardware Implementation of a Hybrid Dynamic Gold Code-Based Countermeasure Against Side-Channel AttacksabstractSide-channel attacks have emerged as the predominant approach for exploiting the weaknesses of cryptographic equipment. Therefore, it is becoming increasingly necessary to prioritize countermeasures that can improve the security level of these implementations. A Mixed-Mode Clock Manager (MMCM) primitive has been utilized in several time-based hiding countermeasures against side-channel attacks. However, they cannot be applied to ASIC implementations because the MMCM is a Xilinx primitive. Consequently, this paper proposes a hybrid dynamic Gold code-based solution to generate multiple different frequencies. The countermeasure combines a pair of preferred polynomials with one ring oscillator, so it is suitable for both FPGA and ASIC designs. The hardware overhead of our suggested architecture is 1.007× and 1.009× in terms of slice LUTs and registers, respectively. The total area cost of the circuit on the CMOS 0.18 um process is 398,835 square micrometers, representing a 1.004x increase compared to the unprotected case. Moreover, the approach is resistant to both standard and sliding window-based Correlation Power Analysis attacks, even when employing UP to one million power traces. Thai-Ha Tran, Duc-Thuan Dam, Binh Kieu-Do-Nguyen, Van-Phuc Hoang, Trong-Thuc Hoang, Cong-Kha Pham |
PST | 2 |
| 2024 | Compacting Side-Channel Measurements With Amplitude Peak Location AlgorithmabstractNowadays, cryptographic algorithms are widely used to build safety mechanisms for specific objects in security services. Nevertheless, these algorithms are implemented in the hardware or software of the physical devices. Consequently, attackers will exploit physical information leakages, such as the device’s power consumption, and use them to get secret keys. The correlation power analysis (CPA) attack is a powerful and efficient cryptographic technique. The evaluation method, however, takes time because many traces are necessary to overcome designs protected by different countermeasures. Therefore, this article proposes a new technique to reduce the computation time by extracting the point of interest (POI) with an interpolation method. The proposal uses the local extreme value and two adjacent samples around it to interpolate the actual peak amplitude. Compared to the conventional CPA, the execution time in our solution is decreased by approximately$9.55\times $, with only 53.32% of the given power traces used for attacking the masking design. Moreover, this technique can deal with the public desynchronized ASCAD database and has better results than recent alignment preprocessing methods. We apply the proposal in the preprocessing step before performing the previously non-profiled deep learning-based attacks. Our suggestion requires only 5000 traces, while the reported attacks fail or require more traces to recover the correct subkey. Thai-Ha Tran, Duc-Thuan Dam, Ba-Anh Dao, Van-Phuc Hoang, Cong-Kha Pham, Trong-Thuc Hoang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |