Trong-Thuc Hoang

dblp:166/7144 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-4078-0836ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 4 first-author · 21 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Countering Side-Channel Attacks With a Dynamic S-Box Based on Affine Transformations and Gold Sequences
abstract
Advanced cryptographic devices employ multiscale countermeasures to bolster resilience against side-channel analysis (SCA). In masking-based defenses, secure substitution-boxes (S-boxes) and effective masking schemes are paramount. Additionally, the time-based hiding techniques, leveraging multiple clocks for individual encryption operations, offer significant protection. This article introduces a novel multiscale countermeasure: an improved tower field masking scheme integrated with an affine transformation-based dynamic S-box. Crucially, we incorporate Gold sequences to generate both a random clock source for horizontal hiding and random values for masking. Extensive evaluation using up to five million power traces demonstrates the robustness of our approach against standard correlation power analysis (CPA) and alignment preprocessing techniques, including sliding window and amplitude peak localization. Experimental results show a measurement-to-disclosure (MTD) improvement of at least$150\times $compared to unprotected implementations using stand-alone masking and$375\times $with our multiscale approach. Furthermore, we demonstrate resilience against recent robust profiled deep learning SCA, which could only recover four subkeys even with one million traces.
Thai-Ha Tran, Duc-Thuan Dam, Tuan-Kiet Dang, Duc-Hung Le, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Compact FALCON FFT/NTT Accelerator for Post-Quantum Cryptography
abstract
FALCON is one of four algorithms selected by NIST to standardize post-quantum cryptography standards. FALCON is a digital signature algorithm based on NTRU lattice with difficulty based on the short vector problem. While Kyber and Dilithium algorithms are only based on NTT operations, FALCON uses both NTT and FFT, which is a barrier to Falcon’s hardware implementation. This paper proposes a compact architecture that supports FFT and NTT for the FALCON algorithm. First, we propose an architecture that executes floating-point and complex number operations with theoretic speed and low area requirements. Then, we design a processing element that performs FFT with complex number operations. Finally, we propose an NTT architecture that reuses the resources used for FFT execution with high parallelism. The FPGA implementation results show that the FFT execution takes 2.048k CCs and 4.608k CCs for the 512-point and 1024-point FFT/IFFT, respectively. The NTT/INTT operation takes 288 CCs for FALCON-512 and 640 CCs for FALCON-1024. The speedup improves from 3× to 9.6× for FFT and up to 18× for NTT implementations compared to previous studies.
Duc-Thuan Dam, Thai-Ha Tran, Trong-Hung Nguyen, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS4
2025 A Low-Latency Polynomial Arithmetic Unit for ML-KEM and ML-DSA Standards
abstract
Existing communication protocols based on public key cryptography (PKC) functions will no longer be secure in the quantum era. NIST has released standards for key encapsulation and digital signature mechanisms based on module lattice (ML-KEM and ML-DSA) to address this challenge. In this paper, we propose a unique, high-performance arithmetic unit capable of performing all the polynomial operations needed for ML-KEM and ML-DSA (KDA). The proposed KDA architecture includes a computational unit that supports one 4×1 NTT configuration for ML-DSA and double 4×1 NTT configurations for ML-KEM. A two-step NTT data flow and configurable memory unit are introduced to reorder and store coefficients for all operations. Moreover, we propose the re-used twiddle factor method for NTT and point-wise multiplication. We have implemented three design versions, ML-KEM standalone, ML-DSA standalone, and KDA, and compared them to the existing studies. The comparison shows that our KDA achieves superior ATP performance, improving ×1.1-×3.6.
Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS4
2025 Live Demonstration: ASIC Implementation of ASCON Lightweight Cryptography for IoT Applications
abstract
The use of IoT devices has increased significantly in recent years, and edge computing in IoT is seen as a new and growing trend in the technology industry. While cryptography is widely used to enhance the security of IoT devices, it also has limitations, such as resource constraints and latency. Lightweight cryptography (LWC) aims to balance resource usage and security while minimizing system costs. Among LWC algorithms, ASCON is a potential target for implementation and cryptoanalysis. A demonstration showcases a system-on-chip (SoC) comprising a RISC-V processor and an ASCON LWC core that implemented the ASCON-128 and ASCON-Hash functions. The SoC was fabricated using a 180nm process.
Khai-Duy Nguyen, Tuan-Kiet Dang, Binh Kieu-Do-Nguyen, Cong-Kha Pham, Trong-Thuc Hoang
ISCAS5
2025 Enhanced Tower Field Mask Scheme with Affine Transformation-based Dynamic S-box
abstract
Masking countermeasures are robust solutions applied to cryptographic devices to improve their side-channel analysis resistance. Implementing substitution boxes (S-boxes) and using an efficient mask scheme are essential for improving security in modern ciphers. Consequently, this paper proposes an improved tower field mask scheme with an affine transformation-based dynamic S-box. The approach resists both Correlation Power Analysis attacks with Hamming Weight and Hamming Distance models, even when employing up to two million power traces. The measurement-to-disclosure improvement for the extracted key byte is at least 158× higher than the previous scheme, while our hardware overhead is around 1.04×. Furthermore, the proposal enhances the devices’ resistance to recent Deep-Learning Side-Channel Analyses.
Thai-Ha Tran, Duc-Thuan Dam, Van-Phuc Hoang, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS4
2025 A Unified Approach to Strong PUF and TRNG Using Ring Generator for Cryptography
abstract
Physical Unclonable Functions (PUFs) and True Random Number Generators (TRNGs) primitives always come in pairs to provide authenticity and unpredictability for cryptographic applications. Specifically, PUF-based authentication presents huge potential for a lightweight, low-power, and efficient solution to secure communication in the Internet of Things (IoT) networks. In any PUF-based scheme, the exchanging materials comprise PUF’s responses and random nonces to generate shared session keys. PUFs offer authentication properties to a device by generating reproducible and device-specific randomness, whereas TRNGs harvest random entropy from physical phenomena to produce completely unpredictable output. This paper introduces a design approach to a unified circuit of PUF and TRNG targeting lightweight and versatile to meet the constrained requirements of IoT devices. The design employs the XOR-Latch (XL) cell to extract uncontrollable manufacturing variances to yield a stable and unique output. Additionally, with specific excitation, it can operate as an oscillator. Multiple XL cells are connected to a ring generator, which serves as a back-end obfuscation structure, to construct a robust strong PUF. Our final design on Xilinx Artix-7 FPGA features a compact hardware footprint of 102 Look-Up Tables (LUTs) and 32 Flip-Flops (FFs), which can be positioned within 26 SLICEs. Various design strategies were employed to assess the feasibility of ASIC implementation. Experimental analyses of the PUF mode performance have shown that the uniformity, uniqueness, and reliability metrics satisfy the standards, and the design is resistant to state-of-the-art modeling attacks. Furthermore, the TRNG function has undergone rigorous testing, including various health checks and standard random tests recommended by the National Institute of Standards and Technology (NIST) and the German Federal Office for Information Security (BSI).
Tuan-Kiet Dang, Khai-Duy Nguyen, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Internet Things J.3
2025 A Timing-Constrained Design Methodology for Radix- 2k NTT in Polynomial Arithmetic
abstract
Polynomial modular multiplication is the most complex and costly operation in homomorphic encryption (HE) and post-quantum cryptography (PQC). Using the Number Theoretic Transform (NTT) helps reduce the complexity of multiplication to quasi-linear O($N\,\textup{log}_{2}N$). Although NTT significantly impacts the performance of HE and PQC, existing NTT-based multipliers often fall short due to inefficient data movement and large memory overhead. Notably, deploying low-latency cryptosystems incurs more significant costs with reduced acceleration gains. To overcome these constraints, we introduce a pioneering methodology called timing-constrained NTT (TCO-NTT). We propose an innovative time-controlled memory (TCM) structure that re-orders and stores coefficients within each stage of the NTT. Then, we employ the divide-and-conquer strategy, allowing freely configurable parallelism levels. Besides, our proposed methodology can generalize to radix-2kNTT and supports any arbitrary polynomial degreeNand scale factorpvalues. We evaluate the proposed TCO-NTT on typical HE and PQC parameter sets across multiple levels of parallelism and radix-2kNTT configurations. FPGA implementation results demonstrate that our TCO-NTT achieves minimal hardware cost while consistently executing the NTT in a near-theoretical execution time. Our area-time product (ATP) reports about LUT-ATP (LATP), FF-ATP (FATP), and BRAM-ATP (BATP) surpass the reported-to-date NTT designs by up to 10.2×, 17.8× and 47.2×. The proposed TCO-NTT sets new records for NTT-based multiplier efficiency, laying the foundation for implementing HE and PQC in real-time applications.
Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Tuan-Kiet Dang, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 Efficient Hardware Implementation of the Lightweight CRYSTALS-Kyber
abstract
Quantum computing raises questions about the security of data encrypted using modern methods. Hence, the National Institute of Standards and Technology (NIST) has undertaken standardization of post-quantum cryptography (PQC) algorithms to defend against attacks from both classical and quantum computers. Following four rounds of evaluation, CRYSTALS-Kyber has been selected for standardization. In this paper, we present an efficient hardware architecture of CRYSTALS-Kyber for resource-constrained IoT devices. Firstly, we propose a compact hash module for CRYSTALS-Kyber. A single buffer is designed to perform padding, hashing, and holding data. Hence, using large FIFOs for data input/output is eliminated. Then, we propose a novel non-memory-based iterative number theoretic transform (NMI-NTT) architecture. Finally, the data flow between modules is optimized to improve parallelization and execution time. Implementation results on an Artix-7 FPGA show that our design consumes minimal hardware resources compared to the designs reported to date, corresponding to 5487 LUTs, 3426 FFs, 1548 SLICEs, 3.5 BRAMs, and 2 DSPs. Our design computes key generation, encapsulation, and decapsulation phases in 3.3/4.5/6.1 K-cycles for Kyber512, 5.6/7.1/9.2 K-cycles for Kyber768, and 8.5/10.1/12.9 K-cycles for Kyber1024, with 185MHz operating frequency. Our area-time-product (ATP) performance outperforms other designs.
Trong-Hung Nguyen, Duc-Thuan Dam, Phuc-Phan Duong, Binh Kieu-Do-Nguyen, Cong-Kha Pham, Trong-Thuc Hoang
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 A Trusted Execution Environment RISC-V System on Chip
abstract
This work proposes a new open-source hardware framework for Trusted Execution Environments (TEEs) on RISC-V systems. The framework is designed to be secure, flexible, and easily upgradable. It includes various cryptographic accelerators and an isolated microcontroller to improve boot performance. The design was implemented and tested on VLSI platforms to demonstrate its feasibility and effectiveness.
Binh Kieu-Do-Nguyen, Khai-Duy Nguyen, Tuan-Kiet Dang, Cong-Kha Pham, Trong-Thuc Hoang
HCS5
2024 RISC-V-Based System-on-Chips for IoT Applications
abstract
The rapidly growing IoT devices pose challenges to power requirements. Traditional power sources, such as batteries, face many limitations, especially regarding durability. By gathering energy from environmental sources, power harvesting promises the future of a fully connected world. Achieving ultralow-voltage operation for direct powering from harvesters involves specific strategies. This necessity gives rise to circuit solutions characterized by low minimum operating voltages, power consumption in the pW range, and resilience against supply fluctuations. This work provides a combined solution to achieve the low-power, low-area target for pure power-harvesting devices: a minimal resource RISC-V processor with ultra-low power, low leakage ASIC technology. We implemented two serial architecture-based RISC-V SoCs, SERV-32I and SERV-32E, on 65-nm SOTB technology. The SERV-32I is a basic implementation of the RISC-V base specification, while the SER-32E implements the embedded specification with 16 registers truncated in the Register File. The lowest power consumption achieved by SERV-32I and SERV-32E is reported at 34 nW and 9.7 nW with a 0.27 V power supply and frequency of 7 kHz and 3 kHz at VDD$=0.27 \text{~V}$, respectively. The SERV-32E processor's footprint is about$28 \%$smaller than the SERV-32I's, while performance only drops by about$5 \%$, with the SERV-32E achieving Dhrystone results of 1.05 DMIPS/MHz and SERV-32I at 1.11 DMIPS/MHz at 50 MHz.
Khai-Duy Nguyen, Tuan-Kiet Dang, Binh Kieu-Do-Nguyen, Cong-Kha Pham, Trong-Thuc Hoang
HCS5
2024 A Strong 4 × 4 S-Box Using an Enhanced Tent Map
abstract
Substitution boxes (S-Boxes) are essential nonlinear components in order to be resistant to the cryptanalytic analysis of modern block ciphers. Given their significance, there is a wide range of S-Box construction techniques. Chaos systems are candidates for constructing S-boxes, with properties characterized by uncertainty, irregularity, and unpredictability. This paper proposes an effective method to create a 4 × 4 S-Box with strong cryptographic properties based on a one-dimensional chaotic map, which is the enhanced tent map. The algorithm used to create S-Boxes uses parameters based on the fundamental characteristics of an S-Box. The cryptanalysis results show that the created S-Box satisfies high-security properties. The output Bit Independence Criterion (BIC) and Strict Avalanche Criterion (SAC) simultaneously reach the ideal value, which no other S-Boxes have ever achieved. Additionally, the other evaluation criteria are comparable to most existing S-Boxes. Owing to these characteristics, the S-Box is appropriate for developing lightweight block ciphers.
Phuc-Phan Duong, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS2
2024 Unified-pipelined NTT Architecture for Polynomial Multiplication in Lattice-based Cryptosystems
abstract
Number Theoretic Transformation (NTT) is commonly employed to speed up polynomial multiplication in post-quantum Lattice-Based Cryptography (LBC). A current trend in NTT hardware design involves using an iterative approach for forward and inverse NTT (INTT) computations. However, this iterative method demands substantial temporary memory and complex memory access patterns. This paper introduces a unified-pipelined NTT architecture for high-performance LBC cryptosystems. Our butterfly units employ a specially crafted Digital Signal Processing (DSP) for modular integer multiplication. Consequently, NTT and INTT calculations are carried out more swiftly with minimal hardware requirements, eliminating the need for DSP and Block Random Access Memory (BRAM). We applied this novel architecture to various parameter sets of LBC and implemented it on the Xilinx FPGA platform for comparison with state-of-the-art studies. Implementation results show that the proposed NTT architectures have outstanding hardware area and operating frequency improvements. The Area Time Product (ATP) is significantly improved, equivalent to at least 53% to 94% compared to the best designs reported to date.
Trong-Hung Nguyen, Nguyen The Binh, Huynh Phuc Nghi, Cong-Kha Pham, Trong-Thuc Hoang
ISCAS5
2024 A Unified OTP and PUF Exploiting Post-Program Current on Standard CMOS Technology
abstract
Root-of-Trust (RoT) uses the hardware primitives to provide security in the integrated electronic systems, preventing attacks against the vital information of the system. The typical hardware primitives are used to generate random numbers and store the system's keys. However, the implementation of these primitives achieves challenges in the system due to the physical phenomena used to obtain the entropy for the random number. On the other hand, the non-volatile memories request special technologies with additional processes and layers. This work presents an implementation of a One-Time Program (OTP) memory using an Anti-Fuse (AF) methodology in 180nm standard CMOS technology. In addition, a Physical Unclonable Function (PUF) is unified, exploiting the post-program current generated in the OTP bit cell. A 128-bit OTP array and high-voltage driver are implemented without any special process and layer, occupying 68000μm2. On the other hand, the PUF implementation achieves 49.02% of uniformity, 49.52% of uniqueness, and 99.88% of reliability in the worst case with Process, Voltage, and Temperature (PVT) variations. In addition, the PUF reports a 2651F2/bit normalized area. The OTP memory application needs 6.7-V to program the AF bit cell.
Ronaldo Serrano, Ckristian Duran, Marco Sarmiento, Khai-Duy Nguyen, Tetsuya Iizuka, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS6
2024 An Efficient Hiding Countermeasure with Xilinx MMCM Primitive in Spread Mode
abstract
The Mixed-Mode Clock Manager (MMCM) is a primitive in Xilinx FPGAs that is designed for generating a wide range of output clock frequencies by utilizing a fixed input clock signal. It has been applied to numerous cryptographic devices to improve their side-channel attack resistance. This paper proposes an efficient hiding countermeasure by using the MMCM in spread spectrum mode. In our suggested architecture, the hardware implementation is given by random dynamic frequency-hopping signals. We could achieve better effectiveness in the occupied bandwidth metrics and found 223 available parameter sets, which is significantly smaller than using 219k distinct sets in a previous study, namely a random dynamic frequency scaling countermeasure. The experimental results indicate that a recent deep learning-based leakage assessment requires nearly one million traces to detect leakage points, whereas the well-known t-test methodology cannot detect any information leakage in five million measurements. Furthermore, this countermeasure is capable of withstanding both conventional and sliding window-based Correlation Power Analysis attacks, despite utilizing up to five million power traces.
Thai-Ha Tran, Van-Phuc Hoang, Duc-Hung Le, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS4
2024 An Efficient Method for Accelerating Kyber and Dilithium Post-Quantum Cryptography
abstract
Post-quantum cryptography (PQC) algorithms were introduced in response to the threats of attacks using quantum computers. The CRYSTALS-Kyber and CRYSTALS-Dilithium are two of the algorithms chosen by NIST to standardize the PQC, which are lattice-based algorithms. Number theoretic transform (NTT) helps lattice-based algorithms reduce latency, but it is still their bottleneck. Along with that, the RISC-V instruction set architecture also opens up flexible methods to solve different problems. This paper proposes a RISC-V system-on-a-chip (SoC) architecture with a computational accelerator for NTT-based calculations for Kyber and Dilithium. Implementation results show that software running on proposed SoC using accelerators has improved in NTT/INTT by up to$36.75\times/42.69\times$compared to software on embedded devices, up to$4.07\times/4.38\times$for software running on RISC-V SoCs, and up to$8.11\times$for NTT of the previous software/hardware architectures.
Duc-Thuan Dam, Trong-Hung Nguyen, Thai-Ha Tran, Binh Kieu-Do-Nguyen, Trong-Thuc Hoang, Cong-Kha Pham
PST5
2024 Hardware Implementation of a Hybrid Dynamic Gold Code-Based Countermeasure Against Side-Channel Attacks
abstract
Side-channel attacks have emerged as the predominant approach for exploiting the weaknesses of cryptographic equipment. Therefore, it is becoming increasingly necessary to prioritize countermeasures that can improve the security level of these implementations. A Mixed-Mode Clock Manager (MMCM) primitive has been utilized in several time-based hiding countermeasures against side-channel attacks. However, they cannot be applied to ASIC implementations because the MMCM is a Xilinx primitive. Consequently, this paper proposes a hybrid dynamic Gold code-based solution to generate multiple different frequencies. The countermeasure combines a pair of preferred polynomials with one ring oscillator, so it is suitable for both FPGA and ASIC designs. The hardware overhead of our suggested architecture is 1.007× and 1.009× in terms of slice LUTs and registers, respectively. The total area cost of the circuit on the CMOS 0.18 um process is 398,835 square micrometers, representing a 1.004x increase compared to the unprotected case. Moreover, the approach is resistant to both standard and sliding window-based Correlation Power Analysis attacks, even when employing UP to one million power traces.
Thai-Ha Tran, Duc-Thuan Dam, Binh Kieu-Do-Nguyen, Van-Phuc Hoang, Trong-Thuc Hoang, Cong-Kha Pham
PST5
2024 Compacting Side-Channel Measurements With Amplitude Peak Location Algorithm
abstract
Nowadays, cryptographic algorithms are widely used to build safety mechanisms for specific objects in security services. Nevertheless, these algorithms are implemented in the hardware or software of the physical devices. Consequently, attackers will exploit physical information leakages, such as the device’s power consumption, and use them to get secret keys. The correlation power analysis (CPA) attack is a powerful and efficient cryptographic technique. The evaluation method, however, takes time because many traces are necessary to overcome designs protected by different countermeasures. Therefore, this article proposes a new technique to reduce the computation time by extracting the point of interest (POI) with an interpolation method. The proposal uses the local extreme value and two adjacent samples around it to interpolate the actual peak amplitude. Compared to the conventional CPA, the execution time in our solution is decreased by approximately$9.55\times $, with only 53.32% of the given power traces used for attacking the masking design. Moreover, this technique can deal with the public desynchronized ASCAD database and has better results than recent alignment preprocessing methods. We apply the proposal in the preprocessing step before performing the previously non-profiled deep learning-based attacks. Our suggestion requires only 5000 traces, while the reported attacks fail or require more traces to recover the correct subkey.
Thai-Ha Tran, Duc-Thuan Dam, Ba-Anh Dao, Van-Phuc Hoang, Cong-Kha Pham, Trong-Thuc Hoang
IEEE Trans. Very Large Scale Integr. Syst.6
2024 Spread Spectrum-Based Countermeasures for Cryptographic RISC-V SoC
abstract
Side-channel analysis attacks have become the primary method for exploiting the vulnerabilities of cryptographic devices. Therefore, focusing on countermeasures to enhance the security level of these implementations evolves even more urgently. This article proposes a time-based hiding countermeasure by using spread-spectrum signals. In our RISC-V system on chip (SoC), cryptographic accelerators are given by random dynamic frequency-hopping signals. We found 223 available parameter sets for a Xilinx Mixed-Mode Clock Manage primitive in spread spectrum mode and achieved better effectiveness in the occupied bandwidth (OBW) metric. The mixed mode clock managers (MMCMs) output signal and the range of frequencies within the spread will be changed randomly, resulting in multiple clocks for individual encryption. The effectiveness of this proposal is demonstrated by conducting realistic side-channel attacks (SCAs) and state-of-the-art leakage assessment methodologies on the well-known data encryption standard, i.e., the Advanced Encryption Standard (AES) accelerator. Even though we used up to five million power traces, the test results show that our defense can stand up to a regular correlation power analysis (CPA) attack as well as alignment preprocessing methods, like CPA attacks that use a sliding window or an amplitude peak location algorithm. Furthermore, the t-test methodology cannot detect any first-order information leakage in five million traces; meanwhile, the deep learning leakage assessment (DLLA) requires nearly one million power traces in the training test to detect leakage points.
Thai-Ha Tran, Ba-Anh Dao, Duc-Hung Le, Van-Phuc Hoang, Trong-Thuc Hoang, Cong-Kha Pham
IEEE Trans. Very Large Scale Integr. Syst.5
2023 Dynamic Gold Code-Based Chaotic Clock for Cryptographic Designs to Counter Power Analysis Attacks
abstract
Research on side-channel attacks has recently made a lot of progress, and one of the most potential solutions is a power analysis attack. Thus, focusing on countermeasures to improve the security level of cryptographic devices is a matter of concern. This paper proposes a time-based hiding countermeasure by using a dynamic Gold code-based chaotic clock. In our work, 36.95% (143 out of 387) of parameter sets that pass two standard statistical test suites can be used to reconfigure the Gold code generator's initial vector. The experimental results demonstrate that our countermeasure helped to harden the targeted Advanced Encryption Standard against several alignment pre-processing methods. When compared to the Random Dynamic Frequency Scaling countermeasure, the number of traces required to reveal the secret key must be increased by three times, but the overhead is approximately four times lower.
Thai-Ha Tran, Anh-Tien Le, Trong-Thuc Hoang, Van-Phuc Hoang, Cong-Kha Pham
ACM Great Lakes Symposium on VLSI3
2023 In-NVRAM Unified PUF and TRNG Based on Standard CMOS Technology
abstract
Hardware security primitives provide Root-of-Trust (RoT) procedures for booting, authentication, and key generation processes in secure integrated systems. The RoT requires True Random Number Generators (TRNGs), Physical Unclonable Functions (PUFs), and non-volatile memories for essential key generation and identity authentication. However, these implementations introduce challenges due to the physical phenomena used in each primitive, requiring complex calibration or special technologies with additional masks. In addition, the integration of separated implementations in a single system-on-a-chip increases the area overhead. This work describes a unified PUF-TRNG in a Non-Volatile Random Access Memory (NVRAM) implementation in 180-nm CMOS technology. The PUF and TRNG primitives are based on the NVRAM metastability in the sense amplifier. The TRNG passes the statistical and entropy tests provided by NIST SP800-22 and SP800-90B, respectively. In addition, the normalized minimum entropy of the TRNG is 0.987 in the worst case with PVT (Process, Voltage, and Temperature) variations. The PUF uniformity, uniqueness and reliability are 49.85%, 48.12% and 99.58%, respectively at nominal conditions. Moreover, the PUF reach$\mathbf{6735} F^{2}/\mathbf{bit}$normalized area11F2= (area)/(minimum feature size of the process)2. The NVRAM needs 8.5V for the programming and erasing modes. Finally, the unified implementation occupies$\mathbf{43155}\mu m^{2}$with$\mathbf{1332}kF^{2}$of normalized area.
Ronaldo Serrano, Marco Sarmiento, Ckristian Duran, Tuan-Kiet Dang, Trong-Thuc Hoang, Cong-Kha Pham
ISCAS5
2023 Transition Factors of Power Consumption Models for CPA Attacks on Cryptographic RISC-V SoC
abstract
Physical cryptographic devices are vulnerable to side-channel information leakages during operation. They are widely used in software as well as hardware implementations, ranging from microcontrollers and microprocessors to hardware accelerators in System on Chips (SoCs). Nowadays, cryptographic RISC-V SoCs are becoming the most prominent solution compared to the rest. Cryptographic accelerators provide users with a very high level of flexibility and customization of chips suited to specific applications in these systems. First, this research aims to confirm the effectiveness of the Correlation Power Analysis attack on cryptographic SoCs based on three different power consumption models. In each model, the effectiveness of an attack depends on the transition factor, which is a ratio related to different characteristics of the device's power consumption. Then, we focus on modifying the configuration on the SoC and attacking the AES hardware implementation on these designs. The experimental results show that applying the Switching Distance model brings the highest performance. With our suggested range of transition factors, the number of traces needed to find the secret key can be reduced by 13.35% in the best case.
Thai-Ha Tran, Ba-Anh Dao, Trong-Thuc Hoang, Van-Phuc Hoang, Cong-Kha Pham
IEEE Trans. Computers3
2022 Spectre attack detection with Neutral Network on RISC-V processor
abstract
The goal of this study is to investigate the problem of caches side-channel attacks on the open-source RISC-V architecture. Traditionally, these concerns have been addressed by research into hardware enhancements or software defense strategies. Those methods are, unfortunately, extremely hard to accomplish or lead to significant performance degradation. In this article, we investigate at a system for identifying cache side-channel threats such Spectre in real time. We utilize performance counters inside the processor to track the processor’s cache behavior and a neural network to evaluate the acquired data. Since the presence of cache side-channels typically results in a dramatically altered cache utilization behaviors, our neural network would take advantage of it and detect a Spectre attack in our test environment with an accuracy of more than 99%. To summarize, we are able to inform the user when a cache side-channel attack occurs using data from only four counters.
Anh-Tien Le, Trong-Thuc Hoang, Ba-Anh Dao, Akira Tsukamoto, Kuniyasu Suzaki, Cong-Kha Pham
ISCAS2
2021 System-on-Chip Implementation of Trusted Execution Environment with Heterogeneous Architecture
abstract
This poster presents a Trusted Execution Environment (TEE) hardware implementation based on a heterogeneous architecture. The TEE verifies the integrity of software applications based on a chain of trust with the initial authentication. The chain-of-trust is implemented in software, using TEE hardware crypto-processors. The initial authentication is called the Root-of-Trust (RoT), and the isolated 32-bit system handles it. On the peripheral bus, there are several cryptography accelerators implemented such as SHA- 3, ED25519, AES, and a True Random Number Generator (TRNG). The TRNG module has not only the public channel over the peripheral bus but also a special private channel just for the isolated core. The proposed system was implemented in a 5mm x 5mm die by the 180-nm ROHM process library.
Trong-Thuc Hoang, Ckristian Duran, Ronaldo Serrano, Marco Sarmiento, Khai-Duy Nguyen, Akira Tsukamoto, Kuniyasu Suzaki, Cong-Kha Pham
HCS1
2021 A CORDIC-based Trigonometric Hardware Accelerator with Custom Instruction in 32-bit RISC-V System-on-Chip
abstract
This poster presents a 32-bit Reduced Instruction Set Computer five (RISC-V) microprocessor with a COordinate Rotation DIgital Computer (CORDIC) algorithm accelerator. The implemented core processor is the VexRiscv CPU, an RV32IM variant of the RISC-V ISA processor. Within the VexRiscv core, the CORDIC accelerator was connected directly to the Execute stage. The core was placed in Briey System-on-Chip (SoC) and was synthesized on Field Programmable Gate Array (FPGA) and on Application Specific Integrated Chip (ASIC) level with the cell logic of ROHM- 180nm technology
Khai-Duy Nguyen, Tuan-Kiet Dang, Trong-Thuc Hoang, Quynh Nguyen Quang Nhu, Cong-Kha Pham
HCS3
2020 Cryptographic Accelerators for Trusted Execution Environment in RISC-V Processors
abstract
The trusted execution environment protects data by taking advantage of memory isolation schemes. Most of the software implementations on security enclaves offer a framework that can be implemented on any processor architecture. Assuming that privilege escalation is not possible through software means, the only way to access protected data is over authentication over a driver in kernel mode. However, the use of hardware back-doors cannot prevent processor execution in more privileged modes. Implementation of kernel-mode allows the reading of sensitive data over the protected regions of memory. In this work, a proposal of crypto-accelerator is described. The peripheral bus in the proposed architecture features a write-only secure memory. That means the cryptography operations on the software level can not read the sensitive data from that secure memory. This approach suppresses any cache coherence manipulator and fault execution-related attacks against reading sensitive data. The peripheral can be useful to accelerate the cryptography operations, and store securely intermediate calculations as well as storing secure keys. The time of execution compared to the software counterpart can be reduced down to 2.5 decades, and the throughput is risen to 3 decades, reaching speeds of 30MB/s for large chunks of data. The total area represents 10.7% of the total area of a dual-core RISC-V processor with RV64IMAFC extensions and TileLink buses.
Trong-Thuc Hoang, Ckristian Duran, Akira Tsukamoto, Kuniyasu Suzaki, Cong-Kha Pham
ISCAS1
2019 Live Demonstration: Real-Time Auto-Exposure Histogram Equalization Video-System using Frequent Items Counter
abstract
In this demonstration, a real-time auto-exposure Histogram Equalization (HE) video-system is presented. The video histogram is extracted in each frame by the Frequent Items Counter (FIC) core. Based on the HE Transformation Function (HE-TF), the camera exposure value is adjusted to fit the current luminance condition. The proposed system was developed on the VEEK-MT-SoCKit with an FPGA chip of Altera Cyclone V SoC and a 5-Megapixel (5-MP) Charge Coupled Device (CCD). The video resolution is 1280×800. The monitor display rate is at 60Hz while the CCD capture rate is at 24.28Hz to 38.98Hz depend on the exposure value. The histogram, the transformation function, and the camera exposure value are changed in each frame to satisfy the real-time requirement.
Takahiro Hosaka, Trong-Thuc Hoang, Van-Phuc Hoang, Duc-Hung Le, Katsumi Inoue, Cong-Kha Pham
ISCAS2
2019 A 1.2-V 90-MHz Bitmap Index Creation Accelerator with 0.27-nW Standby Power on 65-nm Silicon-On-Thin-Box (SOTB) CMOS
abstract
Although bitmap index (BI) can surmount complex and multi-dimensional queries, the creation of BI itself is a time-consuming task. Many studies exploit the highly parallel processing capabilities of multi-core CPUs, graphics processing units (GPUs), or field-programmable gate arrays (FPGAs) to overcome this obstacle. This study, on the other hand, proposes a 65-nm silicon-on-thin-buried-oxide (SOTB) hardware accelerator dedicated to BI creation. The fabricated chip could operate at different supply voltages, from 0.45-V to 1.2-V. Concretely, in the active mode with the supply voltage of 1.2-V, this chip was fully operational at 90-MHz and consumed approximately 88.1-pJ/cycle. In the standby mode with the supply voltage of 0.45-V and clock gated, the power consumption was only 476.1-nW. Moreover, when the reverse back-gate bias voltage of -2.5-V is supplied, the standby power sharply dropped to 0.27-nW or approximately 1,763 times. This achievement is vitally essential for the energy-efficient applications, where the performance should be maximized during peak workload hours and the power should be minimized during off-peak time.
Xuan-Thuan Nguyen, Trong-Thuc Hoang, Katsumi Inoue, Ngoc-Tu Bui, Van-Phuc Hoang, Cong-Kha Pham
ISCAS2
2018 High-speed 8/16/32-point DCT Architecture Using Fixed-rotation Adaptive CORDIC
abstract
In this paper, the high-speed Discrete Cosine Transform (DCT) architecture is presented using the Adaptive CORDIC (ACor) algorithm built with a fixed-rotation angle. The proposed method is implemented in six different versions corresponding to the number of DCT point, i.e., 8-point (8p), 16-point (16p), and 32-point (32p), and the number of ACor stages, i.e., 2-Stage (2S) and 3-Stage (3S). The implementations are built and verified on an Altera Stratix IV FPGA. The 2S designs of 8p-DCT, 16p-DCT, and 32p-DCT achieve the maximum operating frequencies of 179.86 MHz, 162.60 MHz, and 136.97 MHz, respectively. Moreover, the 2S-32p-DCT module is implemented in ASIC with the 65nm-SOTB CMOS technology. The synthesis shows that the core costs 47.2K gates and consumes about 0.68 mW while operating at 100 MHz clock rate. The 2S implementations of 8p-DCT, 16p-DCT, and 32p-DCT achieve four, five, and six adder-delay, mean-square-error of 1.403e-4, 2.029e-2, and 7.663e-2, and coding gain of 8.8108 dB, 9.0984 dB, and 9.2170 dB, respectively. In comparison with recent works, the proposed method achieves the best timing performances, good accuracy results, and adequate resources cost.
Trong-Thuc Hoang, Cong-Kha Pham, Duc-Hung Le
ISCAS1
2018 A 219-μW 1D-to-2D-Based Priority Encoder on 65-nm SOTB CMOS
abstract
Priority encoder (PE) is recognized as an indispensable component in the content-addressable memory. In this paper, two efficient architecture of 64-bit PE and 256-bit PE using 1D-array to 2D-array conversion (1D-to-2D) method are presented and implemented in a 65-nm Silicon-on-thin-buried-oxide (SOTB) CMOS process. The 1D-to-2D method is exploited because of its advantages in large-sized PE construction. The SOTB CMOS process is utilized because of its prominent advantages of low-power and high-performance configuration using back bias voltages. The measurement results at 1.2 V showed that a fabricated PE256 chip was fully operational at 45 MHz and consumed approximately 219 μW. Additionally, in sleep mode, the leakage power dropped as low as 0.34 μW at 0.6 V.
Xuan-Thuan Nguyen, Trong-Thuc Hoang, Hong-Thu Nguyen, Katsumi Inoue, Cong-Kha Pham
ISCAS2
2016 A hybrid adaptive CORDIC in 65nm SOTB CMOS process
abstract
In this paper, a hybird adaptive Coordinate Rotation Digital Computer (HA-CORDIC) has implemented in 65nm Silicon On Thin Buried oxide (SOTB) CMOS technology. In the HA-CORDIC implementation, the adaptive algorithm is utilized for reducing the iteration of CORDIC algorithm. In comparison with other floating-point CORDIC designs, the latency of our proposed scheme is lower. It spends only 12, 20, and 26 clocks cycles in the best, average, and worst case, respectively. The HA-CORDIC exploits some design techniques such as resource sharing, pipeline, and parallel processing to achieve low-resource and low-latency. In 65nm SOTB CMOS technology, this design is able to operate at 50 MHz frequency with 0.5 V supply voltage, 0.36 mA current, and 0.058 mm2 area. Its power consumption of HA-CORDIC is 0.251 mW, about three times lower than the one in conventional CMOS technology. Its leakage current is about 0.492 μA if the supply voltage VDD is 0.4 V and the bias voltage VBB is -1.5 V. This leakage current is about four times lower than that of HA-CORDIC implementing in conventional CMOS.
Trong-Thuc Hoang, Duc-Hung Le, Hong-Thu Nguyen, Xuan-Thuan Nguyen, Cong-Kha Pham
ISCAS1
2016 An efficient FPGA-based database processor for fast database analytics
abstract
Recent years have witnessed a massive growth of global data due to the ubiquitous internet-of-thing products, social networking services, and mobile devices. Fast database analytics, therefore, has been increasingly attractive to numerous research. In this paper, a low-latency FPGA-based Database Processor (DBP) using bitmap index is proposed. By exploiting available embedded memory blocks and logic elements, a 50-MHz DBP is capable of performing 1,024 queries for entire 32,768 4-KB records within around 3.31 ms. In other words, the DBP can analyze a capacity data of nearly 37.76 GB per second.
Xuan-Thuan Nguyen, Hong-Thu Nguyen, Trong-Thuc Hoang, Katsumi Inoue, Osamu Shimojo, Toshio Murayama, Kenji Tominaga, Cong-Kha Pham
ISCAS3