Tuan-Kiet Dang

dblp:282/6966 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-2616-2510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Countering Side-Channel Attacks With a Dynamic S-Box Based on Affine Transformations and Gold Sequences
abstract
Advanced cryptographic devices employ multiscale countermeasures to bolster resilience against side-channel analysis (SCA). In masking-based defenses, secure substitution-boxes (S-boxes) and effective masking schemes are paramount. Additionally, the time-based hiding techniques, leveraging multiple clocks for individual encryption operations, offer significant protection. This article introduces a novel multiscale countermeasure: an improved tower field masking scheme integrated with an affine transformation-based dynamic S-box. Crucially, we incorporate Gold sequences to generate both a random clock source for horizontal hiding and random values for masking. Extensive evaluation using up to five million power traces demonstrates the robustness of our approach against standard correlation power analysis (CPA) and alignment preprocessing techniques, including sliding window and amplitude peak localization. Experimental results show a measurement-to-disclosure (MTD) improvement of at least$150\times $compared to unprotected implementations using stand-alone masking and$375\times $with our multiscale approach. Furthermore, we demonstrate resilience against recent robust profiled deep learning SCA, which could only recover four subkeys even with one million traces.
Thai-Ha Tran, Duc-Thuan Dam, Tuan-Kiet Dang, Duc-Hung Le, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Trans. Very Large Scale Integr. Syst.3
2025 Live Demonstration: ASIC Implementation of ASCON Lightweight Cryptography for IoT Applications
abstract
The use of IoT devices has increased significantly in recent years, and edge computing in IoT is seen as a new and growing trend in the technology industry. While cryptography is widely used to enhance the security of IoT devices, it also has limitations, such as resource constraints and latency. Lightweight cryptography (LWC) aims to balance resource usage and security while minimizing system costs. Among LWC algorithms, ASCON is a potential target for implementation and cryptoanalysis. A demonstration showcases a system-on-chip (SoC) comprising a RISC-V processor and an ASCON LWC core that implemented the ASCON-128 and ASCON-Hash functions. The SoC was fabricated using a 180nm process.
Khai-Duy Nguyen, Tuan-Kiet Dang, Binh Kieu-Do-Nguyen, Cong-Kha Pham, Trong-Thuc Hoang
ISCAS2
2025 A Unified Approach to Strong PUF and TRNG Using Ring Generator for Cryptography
abstract
Physical Unclonable Functions (PUFs) and True Random Number Generators (TRNGs) primitives always come in pairs to provide authenticity and unpredictability for cryptographic applications. Specifically, PUF-based authentication presents huge potential for a lightweight, low-power, and efficient solution to secure communication in the Internet of Things (IoT) networks. In any PUF-based scheme, the exchanging materials comprise PUF’s responses and random nonces to generate shared session keys. PUFs offer authentication properties to a device by generating reproducible and device-specific randomness, whereas TRNGs harvest random entropy from physical phenomena to produce completely unpredictable output. This paper introduces a design approach to a unified circuit of PUF and TRNG targeting lightweight and versatile to meet the constrained requirements of IoT devices. The design employs the XOR-Latch (XL) cell to extract uncontrollable manufacturing variances to yield a stable and unique output. Additionally, with specific excitation, it can operate as an oscillator. Multiple XL cells are connected to a ring generator, which serves as a back-end obfuscation structure, to construct a robust strong PUF. Our final design on Xilinx Artix-7 FPGA features a compact hardware footprint of 102 Look-Up Tables (LUTs) and 32 Flip-Flops (FFs), which can be positioned within 26 SLICEs. Various design strategies were employed to assess the feasibility of ASIC implementation. Experimental analyses of the PUF mode performance have shown that the uniformity, uniqueness, and reliability metrics satisfy the standards, and the design is resistant to state-of-the-art modeling attacks. Furthermore, the TRNG function has undergone rigorous testing, including various health checks and standard random tests recommended by the National Institute of Standards and Technology (NIST) and the German Federal Office for Information Security (BSI).
Tuan-Kiet Dang, Khai-Duy Nguyen, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Internet Things J.1
2025 A Timing-Constrained Design Methodology for Radix- 2k NTT in Polynomial Arithmetic
abstract
Polynomial modular multiplication is the most complex and costly operation in homomorphic encryption (HE) and post-quantum cryptography (PQC). Using the Number Theoretic Transform (NTT) helps reduce the complexity of multiplication to quasi-linear O($N\,\textup{log}_{2}N$). Although NTT significantly impacts the performance of HE and PQC, existing NTT-based multipliers often fall short due to inefficient data movement and large memory overhead. Notably, deploying low-latency cryptosystems incurs more significant costs with reduced acceleration gains. To overcome these constraints, we introduce a pioneering methodology called timing-constrained NTT (TCO-NTT). We propose an innovative time-controlled memory (TCM) structure that re-orders and stores coefficients within each stage of the NTT. Then, we employ the divide-and-conquer strategy, allowing freely configurable parallelism levels. Besides, our proposed methodology can generalize to radix-2kNTT and supports any arbitrary polynomial degreeNand scale factorpvalues. We evaluate the proposed TCO-NTT on typical HE and PQC parameter sets across multiple levels of parallelism and radix-2kNTT configurations. FPGA implementation results demonstrate that our TCO-NTT achieves minimal hardware cost while consistently executing the NTT in a near-theoretical execution time. Our area-time product (ATP) reports about LUT-ATP (LATP), FF-ATP (FATP), and BRAM-ATP (BATP) surpass the reported-to-date NTT designs by up to 10.2×, 17.8× and 47.2×. The proposed TCO-NTT sets new records for NTT-based multiplier efficiency, laying the foundation for implementing HE and PQC in real-time applications.
Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Tuan-Kiet Dang, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 A Trusted Execution Environment RISC-V System on Chip
abstract
This work proposes a new open-source hardware framework for Trusted Execution Environments (TEEs) on RISC-V systems. The framework is designed to be secure, flexible, and easily upgradable. It includes various cryptographic accelerators and an isolated microcontroller to improve boot performance. The design was implemented and tested on VLSI platforms to demonstrate its feasibility and effectiveness.
Binh Kieu-Do-Nguyen, Khai-Duy Nguyen, Tuan-Kiet Dang, Cong-Kha Pham, Trong-Thuc Hoang
HCS3
2024 RISC-V-Based System-on-Chips for IoT Applications
abstract
The rapidly growing IoT devices pose challenges to power requirements. Traditional power sources, such as batteries, face many limitations, especially regarding durability. By gathering energy from environmental sources, power harvesting promises the future of a fully connected world. Achieving ultralow-voltage operation for direct powering from harvesters involves specific strategies. This necessity gives rise to circuit solutions characterized by low minimum operating voltages, power consumption in the pW range, and resilience against supply fluctuations. This work provides a combined solution to achieve the low-power, low-area target for pure power-harvesting devices: a minimal resource RISC-V processor with ultra-low power, low leakage ASIC technology. We implemented two serial architecture-based RISC-V SoCs, SERV-32I and SERV-32E, on 65-nm SOTB technology. The SERV-32I is a basic implementation of the RISC-V base specification, while the SER-32E implements the embedded specification with 16 registers truncated in the Register File. The lowest power consumption achieved by SERV-32I and SERV-32E is reported at 34 nW and 9.7 nW with a 0.27 V power supply and frequency of 7 kHz and 3 kHz at VDD$=0.27 \text{~V}$, respectively. The SERV-32E processor's footprint is about$28 \%$smaller than the SERV-32I's, while performance only drops by about$5 \%$, with the SERV-32E achieving Dhrystone results of 1.05 DMIPS/MHz and SERV-32I at 1.11 DMIPS/MHz at 50 MHz.
Khai-Duy Nguyen, Tuan-Kiet Dang, Binh Kieu-Do-Nguyen, Cong-Kha Pham, Trong-Thuc Hoang
HCS2
2023 In-NVRAM Unified PUF and TRNG Based on Standard CMOS Technology
abstract
Hardware security primitives provide Root-of-Trust (RoT) procedures for booting, authentication, and key generation processes in secure integrated systems. The RoT requires True Random Number Generators (TRNGs), Physical Unclonable Functions (PUFs), and non-volatile memories for essential key generation and identity authentication. However, these implementations introduce challenges due to the physical phenomena used in each primitive, requiring complex calibration or special technologies with additional masks. In addition, the integration of separated implementations in a single system-on-a-chip increases the area overhead. This work describes a unified PUF-TRNG in a Non-Volatile Random Access Memory (NVRAM) implementation in 180-nm CMOS technology. The PUF and TRNG primitives are based on the NVRAM metastability in the sense amplifier. The TRNG passes the statistical and entropy tests provided by NIST SP800-22 and SP800-90B, respectively. In addition, the normalized minimum entropy of the TRNG is 0.987 in the worst case with PVT (Process, Voltage, and Temperature) variations. The PUF uniformity, uniqueness and reliability are 49.85%, 48.12% and 99.58%, respectively at nominal conditions. Moreover, the PUF reach$\mathbf{6735} F^{2}/\mathbf{bit}$normalized area11F2= (area)/(minimum feature size of the process)2. The NVRAM needs 8.5V for the programming and erasing modes. Finally, the unified implementation occupies$\mathbf{43155}\mu m^{2}$with$\mathbf{1332}kF^{2}$of normalized area.
Ronaldo Serrano, Marco Sarmiento, Ckristian Duran, Tuan-Kiet Dang, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS4
2021 A CORDIC-based Trigonometric Hardware Accelerator with Custom Instruction in 32-bit RISC-V System-on-Chip
abstract
This poster presents a 32-bit Reduced Instruction Set Computer five (RISC-V) microprocessor with a COordinate Rotation DIgital Computer (CORDIC) algorithm accelerator. The implemented core processor is the VexRiscv CPU, an RV32IM variant of the RISC-V ISA processor. Within the VexRiscv core, the CORDIC accelerator was connected directly to the Execute stage. The core was placed in Briey System-on-Chip (SoC) and was synthesized on Field Programmable Gate Array (FPGA) and on Application Specific Integrated Chip (ASIC) level with the cell logic of ROHM- 180nm technology
Khai-Duy Nguyen, Tuan-Kiet Dang, Trong-Thuc Hoang, Quynh Nguyen Quang Nhu, Cong-Kha Pham
HCS2