EDBT 2026 Demo / reviewers in the wild / expert
Wangchen Dai
dblp:194/8696
· DBLP profile ↗
16ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-5192-1649ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 7 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Less is More: Latent Diffusion for Efficient IoT Side-Channel AnalysisabstractThe proliferation of cryptographic primitives in resource-constrained Internet of Things (IoT) devices has made them prime targets for Side-Channel Analysis (SCA). However, designing effective defenses against these attacks has become increasingly complex, as traditional deep learning approaches rely heavily on extensive profiling datasets that are difficult to obtain in the context of widely distributed and physically restricted IoT environments. This challenge is further exacerbated by countermeasures such as clock jitter and random delays. To overcome this limitation, this paper introduces a novel and data-efficient three-stage framework for generating high-fidelity synthetic side-channel traces. First, we employ a Supervised Variational Autoencoder (S-VAE) to map noisy, high-dimensional raw traces into a compact and denoised latent space, effectively creating an information-rich manifold. Second, a conditional Denoising Diffusion Implicit Model (DDIM), powered by an advanced attention-augmented U-Net, is trained exclusively on this computationally tractable latent space to learn the complex conditional data distribution. Finally, we empirically validate our framework on the public ASCAD benchmark and ChipWhisperer CW308T UFO platform. The results are compelling: an attack model trained solely on our synthetic data successfully recovers the secret key in the most challenging ASCAD_desync100 scenario using only 4107 traces and using only 2560 traces, 97.7% accuracy can be achieved on the Chipwhisphere platform. This work provides a practical and efficient pathway for the robust security evaluation of cryptographic implementations in data-scarce IoT environments, significantly lowering the barrier for thorough side-channel vulnerability analysis. Donald Donglong Chen, Wangchen Dai, Jinfa Hong, Yu Hin Chan, Çetin Kaya Koç, Patrick S. Y. Hung, Ray C. C. Cheung |
IEEE Internet Things J. | 3 |
| 2026 | HI-CKKS: Is High-Throughput Neglected? Reimagining CKKS Efficiency With ParallelismabstractThe rapid advancement of the Industrial Internet of Things (IIoT) has positioned data privacy protection as a critical challenge in smart manufacturing. The Cheon–Kim–Kim–Song (CKKS) homomorphic encryption scheme, with its floating-point approximation and single instruction multiple data support capabilities, has emerged as a vital technology for IIoT privacy preservation. However, the high computational complexity of homomorphic multiplication significantly constrains its application performance in real-world industrial scenarios. While existing optimizations primarily focus on latency reduction, research on high-throughput parallel optimization remains notably insufficient. To address this, we propose HI-CKKS, a GPU-based high-throughput CKKS homomorphic multiplication optimization scheme for IIoT, achieving performance breakthroughs through multilevel technical innovations. By designing a batch-processing asynchronous execution architecture, we resolve host-server interaction delays in edge-cloud collaboration. A hierarchical hybrid optimization strategy combining instruction set acceleration, kernel fusion, and dynamic memory management significantly enhances (I)NTT operation efficiency. In addition, we construct a multidimensional parallel homomorphic multiplication model tailored for IIoT high-concurrency task characteristics. Experimental results show that HI-CKKS improves the throughput of NTT, INTT, and homomorphic multiplication by$175.08\times$,$191.27\times$, and$679.57\times$over CPU baselines. When compared to state-of-the-art GPU implementations on similar platforms, HI-CKKS attains$1.54\times$–$14.70\times$speedup for (I)NTT and$55.8\times$–$266.87\times$speedup for homomorphic multiplication. Our work provides an efficient solution for secure computation in IIoT and offers a new approach to performance optimization of homomorphic encryption in edge-cloud collaborative scenarios. Fuyuan Chen, Jiankuo Dong, Zhenjiang Dong, Wangchen Dai |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration
Hanyu Wei, Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Yunlei Zhao |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | ECO-CRYSTALS: Efficient Cryptography CRYSTALS on Standard RISC-V ISAabstractThe field of post-quantum cryptography (PQC) is continuously evolving. Many researchers are exploring efficient PQC implementation on various platforms, including x86, ARM, FPGA, GPU, etc. In this paper, we present an Efficient CryptOgraphy CRYSTALS (ECO-CRYSTALS) implementation on standard 64-bit RISC-V Instruction Set Architecture (ISA). The target schemes are two winners of the National Institute of Standards and Technology (NIST) PQC competition: CRYSTALS-Kyber and CRYSTALS-Dilithium, where the two most time-consuming operations are Keccak and polynomial multiplication. Notably, this paper is the first highly-optimized assembly software implementation to deploy Kyber and Dilithium on the 64-bit RISC-V ISA. Firstly, we propose a better scheduling strategy for Keccak, which is specifically tailored for the 64-bit dual-issue RISC-V architecture. Our 24-round Keccak permutation (Keccak-$p$[1600,24]) achieves a 59.18% speed-up compared to the reference implementation. Secondly, we apply two modular arithmetic (Montgomery arithmetic and Plantard arithmetic) in the polynomial multiplication of Kyber and Dilithium to get a better lazy reduction. Then, we propose a flexible dual-instruction-issue scheme of Number Theoretic Transform (NTT). As for the matrix-vector multiplication, we introduce a row-to-column processing methodology to minimize the expensive memory access operations. Compared to the reference implementation, we obtain a speedup of 53.85%$\thicksim$85.57% for NTT, matrix-vector multiplication, and INTT in our ECO-CRYSTALS. Finally, the ECO-CRYSTALS implementation for key generation, encapsulation, and decapsulation in Kyber achieves 399k, 448k, and 479k cycles respectively, achieving speedups of 60.82%, 63.93%, and 65.56% compared to the NIST reference implementation. Similarly, the ECO-CRYSTALS implementation for key generation, sign, and verify in Dilithium reaches 1 364k, 3 191k, and 1 369k cycles, showcasing speedups of 54.84%, 64.98%, and 57.20%, respectively. Xinyi Ji, Jiankuo Dong, Junhao Huang 0001, Zhijian Yuan, Wangchen Dai, Fu Xiao 0001, Jingqiang Lin 0001 |
IEEE Trans. Computers | 5 |
| 2024 | REALISE-IoT: RISC-V-Based Efficient and Lightweight Public-Key System for IoT ApplicationsabstractLoRa is a promising choice for deploying an IoT network due to its lightweight feature and the extensive support by LoRa Alliance. However, as a fundamental part of LoRa, the typical LoRaWAN protocol confronts severe security challenges because it insecurely utilizes AES-128 to support the low-cost feature. In this paper, we propose a systematic solution that is compatible with LoRaWAN for IoT applications. We extend the standard LoRaWAN protocol with public-key infrastructures. Public-key features like Key exchange and authentication are supported by lightweight hardware implementations of SHA-2, ECDH, EdDSA, and TRNG. A lightweight RISC-V processor with a security coprocessor is implemented and verified using FPGA technology. The security protocol and the prototype hardware system are validated and evaluated on practical applications from our industrial partner. The prototyped development board consumes a static power of 0.116Wand a dynamic power of 0.206 W. The proposed system can achieve a 5.6x-144.7x speed up and reduce memory usage by 2.4x-12.3x for security computations. Gaoyu Mao, Yao Liu 0006, Wangchen Dai, Guangyan Li, Alan H. F. Lam, Ray C. C. Cheung |
IEEE Internet Things J. | 3 |
| 2024 | Efficiency Optimization Techniques in Privacy-Preserving Federated Learning With Homomorphic Encryption: A Brief SurveyabstractFederated learning (FL) offers distributed machine learning on edge devices. However, the FL model raises privacy concerns. Various techniques, such as homomorphic encryption (HE), differential privacy, and multiparty cooperation, are used to address the privacy issues of the FL model. Among them, HE ensures greater security and privacy since end-to-end encryption maintains data privacy throughout the computation process. Compared with other privacy-preserving techniques, HE does not require the establishment of a trusted environment or protocol among multiple parties and does not involve any artificial noise that can impair system performance. Unfortunately, it suffers from efficiency overhead when applied to privacy-preserving FL (PPFL). Some existing surveys on PPFL discuss the generic construction and organization of PPFL from the perspective of practical HE deployment in PPFL. However, none of them covers the efficiency optimization of HE when applied to PPFL. This article conducts a comprehensive review of the efficiency optimization of HE when applied to PPFL. First, we review general optimization strategies and discuss their limitations when applied directly to HE-based PPFL. Second, an overview of algorithmic, hardware, and hybrid optimizations is provided, along with a discussion of their adaptation. Additionally, we provide a detailed taxonomy of optimizations. Finally, we suggest future HE-based PPFL research directions. Qipeng Xie, Siyang Jiang, Linshan Jiang, Yongzhi Huang 0002, Salabat Khan, Wangchen Dai, Zhe Liu 0001, Kaishun Wu |
IEEE Internet Things J. | 7 |
| 2024 | Leveraging GPU in Homomorphic Encryption: Framework Design and Analysis of BFV VariantsabstractHomomorphic Encryption (HE) enhances data security by enabling computations on encrypted data, advancing privacy-focused computations. The BFV scheme, a promising HE scheme, raises considerable performance challenges. Graphics Processing Units (GPUs), with considerable parallel processing abilities, offer an effective solution. In this work, we present an in-depth study on accelerating and comparing BFV variants on GPUs, including Bajard-Eynard-Hasan-Zucca (BEHZ), Halevi-Polyakov-Shoup (HPS), and recent variants. We introduce a universal framework for all variants, propose optimized BEHZ implementation, and first support HPS variants with large parameter sets on GPUs. We also optimize low-level arithmetic and high-level operations, minimizing instructions for modular operations, enhancing hardware utilization for base conversion, and implementing efficient reuse strategies and fusion methods to reduce computational and memory consumption. Leveraging our framework, we offer comprehensive comparative analyses. Performance evaluation shows a 31.9$\times$speedup over OpenFHE running on a multi-threaded CPU and 39.7% and 29.9% improvement for tensoring and relinearization over the state-of-the-art GPU BEHZ implementation. The leveled HPS variant records up to 4$\times$speedup over other variants, positioning it as a highly promising alternative for specific applications. Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Lu Zhou 0002, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Computers | 3 |
| 2024 | Phantom: A CUDA-Accelerated Word-Wise Homomorphic Encryption LibraryabstractHomomorphic encryption (HE) is a promising technique for privacy-preserving computations, especially the word-wise HE schemes that allow batching. However, the high computational overhead hinders the deployment of HE in real-word applications. GPUs are often used to accelerate execution, but a comprehensive performance comparison of different schemes on the same platform is still missing. In this work, we fill this gap by implementing three word-wise HE schemes BGV, BFV, and CKKS on GPU, with both theoretical and engineering optimizations. We enhance the hybrid key-switching technique, significantly reducing the computational and memory overhead. We explore several kernel fusing strategies to reuse data, resulting in reduced memory access and IO latency, and enhancing the overall performance. By comparing with the state-of-the-art works, we demonstrate the effectiveness of our implementation. Meanwhile, we introduce a unified framework that finely integrates our implementation of the three schemes, covering almost all scheme functions and homomorphic operations. We optimize the management of pre-computation, RNS bases, and memory in the framework, to provide efficient and low-latency data access and transfer. Based on this framework, we provide a thorough benchmark of the three schemes, which can serve as a reference for scheme selection and implementation in constructing privacy-preserving applications. Hao Yang 0062, Shiyu Shen 0001, Wangchen Dai, Lu Zhou 0002, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | Yet Another Improvement of Plantard Arithmetic for Faster Kyber on Low-End 32-bit IoT DevicesabstractIn 2022, the National Institute of Standards and Technology (NIST) made an announcement regarding the standardization of Post-Quantum Cryptography (PQC) candidates. Out of all the Key Encapsulation Mechanism (KEM) schemes, the CRYSTAL-Kyber emerged as the sole winner. This paper presents another improved version of Plantard arithmetic that could speed up Kyber implementations on two low-end 32-bit IoT platforms (ARM Cortex-M3 and RISC-V) without SIMD extensions. Specifically, we further enlarge the input range of the Plantard arithmetic without modifying its computation steps. After tailoring the Plantard arithmetic for Kyber’s modulus, we show that the input range of the Plantard multiplication by a constant is at least 2.14× larger than the original design in TCHES2022. Then, two optimization techniques for efficient Plantard arithmetic on Cortex-M3 and RISC-V are presented.We show that the Plantard arithmetic supersedes both Montgomery and Barrett arithmetic on low-end 32-bit platforms. With the enlarged input range and the efficient implementation of the Plantard arithmetic on these platforms, we propose various optimization strategies for NTT/INTT. We minimize or entirely eliminate the modular reduction of coefficients in NTT/INTT by taking advantage of the larger input range of the proposed Plantard arithmetic on low-end 32-bit platforms. Furthermore, we propose two memory optimization strategies that reduce 23.50%~28.31% stack usage for the speed-version Kyber implementation when compared to its counterpart on Cortex-M4. The proposed optimizations make the speed-version implementation more feasible on low-end IoT devices. Thanks to the aforementioned optimizations, our NTT/INTT implementation shows considerable speedups compared to the state-of-the-art work. Overall, we demonstrate the applicability of the speed-version Kyber implementation on memory-constrained IoT platforms and set new speed records for Kyber on these platforms. Junhao Huang 0001, Haosong Zhao, Jipeng Zhang 0001, Wangchen Dai, Lu Zhou 0002, Ray C. C. Cheung, Çetin Kaya Koç, Donald Donglong Chen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | High-Throughput GPU Implementation of Dilithium Post-Quantum Digital SignatureabstractDigital signatures are fundamental building blocks in various protocols to provide integrity and authenticity. The development of the quantum computing has raised concerns about the security guarantees afforded by classical signature schemes. CRYSTALS-Dilithium is an efficient post-quantum digital signature scheme based on lattice cryptography and has been selected as the primary algorithm for standardization by the National Institute of Standards and Technology. In this work, we present a high-throughput GPU implementation of Dilithium. For individual operations, we employ a range of computational and memory optimizations to overcome sequential constraints, reduce memory usage and IO latency, address bank conflicts, and mitigate pipeline stalls. This results in high and balanced compute throughput and memory throughput for each operation. In terms of concurrent task processing, we leverage task-level batching to fully utilize parallelism and implement a memory pool mechanism for rapid memory access. We propose a dynamic task scheduling mechanism to improve multiprocessor occupancy and significantly reduce execution time. Furthermore, we apply asynchronous computing and launch multiple streams to hide data transfer latencies and maximize the computing capabilities of both CPU and GPU. Across all three security levels, our GPU implementation achieves over 160× speedups for signing and over 80× speedups for verification on both commercial and server-grade GPUs. This achieves microsecond-level amortized execution times for each task, offering a high-throughput and quantum-resistant solution suitable for a wide array of applications in real systems. Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | ProgramGalois: A Programmable Generator of Radix-4 Discrete Galois Transformation Architecture for Lattice-Based CryptographyabstractLattice-based cryptography (LBC) has been established as a prominent research field, with particular attention on post-quantum cryptography (PQC) and fully homomorphic encryption (FHE). As the implementing bottleneck of PQC and FHE, number theoretic transform (NTT) has been extensively studied. However, current works struggled with scalability, hindering their adaptation to various parameters, such as bit width and polynomial length. In this article, we proposed a novel Discrete Galois Transformation (DGT) algorithm utilizing the radix-4 variant to achieve a higher level of parallelism to the existing NTT. Furthermore, to implement the efficient radix-4 DGT adapting more LBCs, we proposed a set of scalable building blocks, including a modified Barrett modular multiplier accepting arbitrary modulus with only one integer multiplier, a radix-4 DGT butterfly unit, and a stream permutation network. The proposed modules are implemented on the Xilinx Virtex-7 and U250 FPGA to evaluate resource utilization and performance. Lastly, a design space exploration framework is proposed to generate optimized radix-4 DGT hardware constrained by polynomial and platform parameters. The sensitivity analysis showcases the generated hardware’s performance and scalability. The implementation results on the Xilinx Virtex-7 and U250 FPGA show significant performance improvements over the state-of-the-art works, which reached at least 35%, 192%, and 68% area-time product improvements in terms of LUTs, BRAMs, and DSPs, respectively. Guangyan Li, Zewen Ye, Donald Donglong Chen, Wangchen Dai, Gaoyu Mao, Kejie Huang, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2023 | IoT-based generalized multi-granulation sequential three-way decisions
Yongjing Zhang, Wangchen Dai, Chengxin Hong |
Comput. Commun. | 3 |
| 2023 | High-performance and Configurable SW/HW Co-design of Post-quantum Signature CRYSTALS-DilithiumabstractCRYSTALS-Dilithium is a lattice-based post-quantum digital signature scheme that is resistant to attacks by quantum computers and has been selected to be standardized in the NIST post-quantum cryptography (PQC) standardization process. However, the speed performance and design flexibility of the Dilithium still need to be evaluated. This article presents a high-performance software/hardware co-design of CRYSTALS-Dilithium based on the NIST PQC round-3 parameters. High-speed pipelined hardware modules for NTT/INTT, point-wise multiplication/addition, and for SHAKE are included in the design to accelerate the time-consuming operations in Dilithium. All hardware modules are parameterized, thus allowing full support of runtime configuration to increase versatility. Moreover, the proposed software/hardware architecture and tight operating workflows reduce the data transmission overhead between the processor and other hardware modules. The hardware accelerator is implemented with a reconfigurable logic on FPGA and is integrated with the high-performance ARM Cortex-A9 processor in the Xilinx Zynq Architecture. We measure the performance of the software/hardware system for Dilithium in NIST security levels 2, 3, and 5. Compared to pure software implementations, we achieve 8.7–12.5 times speedup in Key generation, 6.3–7.3 times speedup in Sign, and 9.1–12.2 times speedup in Verify operations. Gaoyu Mao, Donald Donglong Chen, Guangyan Li, Wangchen Dai, Abdurrashid Ibrahim Sanka, Çetin Kaya Koç, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2021 | Scalable Fully Pipelined Hardware Architecture for In-Network Aggregated AllReduce CommunicationabstractThe Ring-AllReduce framework is currently the most popular solution to deploy industry-level distributed machine learning tasks. However, only about half of the maximum bandwidth can be achieved in the optimal condition. In recent years, several in-network aggregation frameworks have been proposed to overcome the drawback, but limited hardware information have been disclosed. In this paper, we propose a scalable fully-pipelined architecture that handles tasks like forwarding, aggregation and retransmission with no bandwidth loss. The architecture is implemented on a Xilinx Ultrascale FPGA that connects to 8 working servers with 10 Gb/s network adapters, and it is able to scale to more complicated scenarios involving more workers. Compared with Ring-AllReduce, using AllReduce-Switch improves the efficient bandwidth of AllReduce communication with a ratio of$1.75\times $. In image training tasks, the proposed hardware architecture helps to achieve up to$1.67\times $speedup to the training process. For computing-intensive models, the speedup from communication may be partially hidden by computing. In particular, for ResNet-50, AllReduce-Switch improves the training process with MPI and NCCL by$1.30\times $and$1.04\times $respectively. Yao Liu 0006, Junyi Zhang 0005, Shuo Liu 0002, Qiaoling Wang, Wangchen Dai, Ray C. C. Cheung |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2018 | FFT-Based McLaughlin's Montgomery Exponentiation without Conditional SelectionsabstractModular multiplication forms the basis of many cryptographic functions such as RSA, Diffie-Hellman key exchange, and ElGamal encryption. For large RSA moduli, combining the fast Fourier transform (FFT) with McLaughlin's Montgomery modular multiplication (MLM) has been validated to offer cost-effective implementation results. However, the conditional selections in McLaughlin's algorithm are considered to be inefficient and vulnerable to timing attacks, since extra long additions or subtractions may take place and the running time of MLM varies. In this work, we restrict the parameters of MLM by a set of new bounds and present a modified MLM algorithm involving no conditional selection. Compared to the original MLM algorithm, we inhibit extra operations caused by the conditional selections and accomplish constant running time for modular multiplications with different inputs. As a result, we improve both area-time efficiency and security against timing attacks. Based on the proposed algorithm, efficient FFT-based modular multiplication and exponentiation are derived. Exponentiation architectures with dual FFT-based multipliers are designed obtaining area-latency efficient solutions. The results show that our work offers a better efficiency compared to the state-of-the-art works from and above 2048-bit operand sizes. For single FFT-based modular multiplication, we have achieved constant running time and obtained area-latency efficiency improvements up to 24.3 percent for 1,024-bit and 35.5 percent for 4,096-bit operands, respectively. Wangchen Dai, Donald Donglong Chen, Ray C. C. Cheung, Çetin Kaya Koç |
IEEE Trans. Computers | 1 |
| 2017 | Area-Time Efficient Architecture of FFT-Based Montgomery MultiplicationabstractThe modular multiplication operation is the most time-consuming operation for number-theoretic cryptographic algorithms involving large integers, such as RSA and Diffie-Hellman. Implementations reveal that more than 75 percent of the time is spent in the modular multiplication function within the RSA for more than 1,024-bit moduli. There are fast multiplier architectures to minimize the delay and increase the throughput using parallelism and pipelining. However such designs are large in terms of area and low in efficiency. In this paper, we integrate the fast Fourier transform (FFT) method into the McLaughlin's framework, and present an improved FFT-based Montgomery modular multiplication (MMM) algorithm achieving high area-time efficiency. Compared to the previous FFT-based designs, we inhibit the zero-padding operation by computing the modular multiplication steps directly using cyclic and nega-cyclic convolutions. Thus, we reduce the convolution length by half. Furthermore, supported by the number-theoretic weighted transform, the FFT algorithm is used to provide fast convolution computation. We also introduce a general method for efficient parameter selection for the proposed algorithm. Architectures with single and double butterfly structures are designed obtaining low area-latency solutions, which we implemented on Xilinx Virtex-6 FPGAs. The results show that our work offers a better area-latency efficiency compared to the state-of-the-art FFT-based MMM architectures from and above 1,024-bit operand sizes. We have obtained area-latency efficiency improvements up to 50.9 percent for 1,024-bit, 41.9 percent for 2,048-bit, 37.8 percent for 4,096-bit and 103.2 percent for 7,680-bit operands. Furthermore, the operating latency is also outperformed with high clock frequency for length-64 transform and above. Wangchen Dai, Donald Donglong Chen, Ray C. C. Cheung, Çetin Kaya Koç |
IEEE Trans. Computers | 1 |