EDBT 2026 Demo / reviewers in the wild / expert
Phap Duong-Ngoc
dblp:268/1120
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-0311-9387ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Number Theoretic Transform Architecture for Homomorphic EncryptionabstractFully homomorphic encryption (FHE) is an innovative cryptographic technology that has the potential to protect the privacy and confidentiality of data in the untrusted environments, such as public clouds or external parties. However, due to the inclusion of time-consuming polynomial arithmetic, FHE remains a challenge for computationally heavy applications. The number theoretic transform (NTT) is widely used in HE to reduce the complexity of polynomial multiplication. Therefore, implementing NTT in hardware for FHE has been explored in prior studies. However, due to the high hardware resource requirements, especially with a large number of moduli, hardware architecture supporting both NTT and its inverse transform (INTT) is still missing. This brief presents a hardware architecture for$2^{17}$NTT and INTT suitable for high-circuit depth CKKS-based HE schemes, satisfying both criteria of high speed and affordability for various FPGA platforms. The implementation results highlight that this design is area-efficient compared to the most related work and hardware-friendly for practical HE-based applications on FPGA devices. Quang Dang Truong, Phap Duong-Ngoc, Hanho Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Area-Efficient Number Theoretic Transform Architecture for Homomorphic EncryptionabstractHomomorphic encryption (HE) has emerged as an ideal cryptographic technology for meaningful computations on encrypted data. Not only does HE secure private information even if the ciphertext is leaked, but it also maintains data integrity when inferring cloud-side services. However, homomorphic computations include expensive polynomial arithmetic, especially polynomial multiplication. Prior studies proposed number theoretic transform (NTT) hardware designs to accelerate polynomial multiplication. However, the trade-off between hardware complexity and throughput of NTT designs was not considered carefully. This paper proposes an area-efficient NTT architecture suitable for HE schemes. Center of the proposed NTT architecture is a high-throughput butterfly unit array, which communicates with a single data memory unit through a conflict-free memory access pattern. Additionally, we developed a twiddle factor generator to reduce memory consumption. The proposed NTT architecture was successfully accelerated on the Xilinx FPGA devices. Performing with a large number of moduli, the proposed NTT design achieves higher hardware efficiency than the prior arts. Especially, our NTT design consumes less on-chip memory with efficiency improvement of$8.8\times $over the most related work. The implementation results confirm that our design methodology has advantages to deploy many NTT accelerators on an FPGA device for practical HE-based applications. Phap Duong-Ngoc, Sunmin Kwon, Donghoon Yoo, Hanho Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | An Efficient Unified Polynomial Arithmetic Unit for CRYSTALS-DilithiumabstractThe CRYSTALS-Dilithium protocol is considered as one of the most promising digital signature schemes in NIST’s post-quantum cryptography standardization process. While separating arithmetic computation units can be advantageous in some cases, it can lead to increased hardware resource consumption and performance degradation. To overcome this issue, this paper proposes a novel architecture called the Unified Polynomial Arithmetic Unit (UniPAU), specifically designed for the Dilithium signature scheme. The proposed UniPAU offers a unique hardware module that can execute all the polynomial operations required for the Dilithium signature scheme. To demonstrate the effectiveness of our design, we implemented it on the Xilinx Zynq UltraScale+ ZCU102 (xczu9eg-ffvb1156-2-e) FPGA platform and evaluated its hardware efficiency and performance. Our implementation results indicate that the proposed UniPAU can achieve comparable throughput while consuming fewer hardware resources compared to state-of-the-art studies. These findings suggest that our UniPAU can provide an optimized and efficient hardware solution for polynomial arithmetic operations in the Dilithium signature scheme. Thang Xuan Pham, Phap Duong-Ngoc, Hanho Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |