VLDB 2026 Research / reviewers in the wild / expert
Donald Donglong Chen
dblp:124/7193 · also Donglong Chen
· DBLP profile ↗
24ranked-venue papers
2as first author
19since 2021 · last 2027
0000-0001-5357-7442ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Security and privacy · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Layer-coupled co-evolutionary cellular automata: A cross-plane diffusion approach for chaotic image privacy protection
Chi Duan, Pengbo Liu 0001, Yong Zhang 0018, Wenjun Zeng 0002, Donald Donglong Chen |
Expert Syst. Appl. | 6 |
| 2026 | Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model AgentsabstractState-of-the-art single-agent claim verification methods struggle with complex claims that require nuanced analysis of multifaceted evidence. Inspired by real-world professional fact-checkers, we propose DebateCV, the first debate-driven claim verification framework powered by multiple LLM agents. In DebateCV, two Debaters argue opposing stances to surface subtle errors in single-agent assessments. A decisive Moderator is then required to weigh the evidential strength of conflicting arguments to deliver an accurate verdict. Yet, zero-shot Moderators are biased toward neutral judgments, and no datasets exist for training them. To bridge this gap, we propose Debate-SFT, a post-training framework that leverages synthetic data to enhance agents' ability to effectively adjudicate debates for claim verification. Results show that our methods surpass state-of-the-art non-debate approaches in both accuracy (across various evidence conditions) and justification quality. Haorui He, Yupeng Li 0001, Dacheng Wen, Yang Chen 0001, Reynold Cheng, Donald Donglong Chen, Francis C. M. Lau 0001 |
WWW | 6 |
| 2026 | Less is More: Latent Diffusion for Efficient IoT Side-Channel AnalysisabstractThe proliferation of cryptographic primitives in resource-constrained Internet of Things (IoT) devices has made them prime targets for Side-Channel Analysis (SCA). However, designing effective defenses against these attacks has become increasingly complex, as traditional deep learning approaches rely heavily on extensive profiling datasets that are difficult to obtain in the context of widely distributed and physically restricted IoT environments. This challenge is further exacerbated by countermeasures such as clock jitter and random delays. To overcome this limitation, this paper introduces a novel and data-efficient three-stage framework for generating high-fidelity synthetic side-channel traces. First, we employ a Supervised Variational Autoencoder (S-VAE) to map noisy, high-dimensional raw traces into a compact and denoised latent space, effectively creating an information-rich manifold. Second, a conditional Denoising Diffusion Implicit Model (DDIM), powered by an advanced attention-augmented U-Net, is trained exclusively on this computationally tractable latent space to learn the complex conditional data distribution. Finally, we empirically validate our framework on the public ASCAD benchmark and ChipWhisperer CW308T UFO platform. The results are compelling: an attack model trained solely on our synthetic data successfully recovers the secret key in the most challenging ASCAD_desync100 scenario using only 4107 traces and using only 2560 traces, 97.7% accuracy can be achieved on the Chipwhisphere platform. This work provides a practical and efficient pathway for the robust security evaluation of cryptographic implementations in data-scarce IoT environments, significantly lowering the barrier for thorough side-channel vulnerability analysis. Donald Donglong Chen, Wangchen Dai, Jinfa Hong, Yu Hin Chan, Çetin Kaya Koç, Patrick S. Y. Hung, Ray C. C. Cheung |
IEEE Internet Things J. | 2 |
| 2026 | Less is more: Clustering with adaptive probability for heterogeneous federated learning
Yunan Wei, Donald Donglong Chen, Minghao Zhao 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Cryptanalysis and Improvement of a Video Cryptosystem via Chaos and S-BoxabstractIn recent years, chaos-based multimedia cryptosystems have gained prominence due to their nonlinear properties, such as sensitivity to initial conditions and long-term unpredictability, which parallel the cryptographic requirements of confusion and diffusion. However, many such systems lack standardization and deviate from secure design principles, resulting in practical vulnerabilities. This article presents a comprehensive cryptanalysis of a widely cited video cryptosystem. The target system combines S-box substitution—derived from either a 12D chaotic map or the Ikeda delay differential equation (DDE)—with a Cipher Block Chaining (CBC) diffusion scheme. Through an analysis of the encryption structure, fundamental design defects are identified. Specifically, the CBC diffusion mechanism lacks key dependence, and the system exhibits an excessive reliance on fixed, publicly exposed S-boxes. These vulnerabilities are demonstrated through chosen-plaintext attacks, known-plaintext attacks, and chosen-ciphertext attacks. Experimental results confirm that the S-box can be fully reconstructed, thereby facilitating the complete decryption of video content in the absence of the secret key. Beyond exposing vulnerabilities, this study offers constructive contributions by proposing six concrete improvement strategies, including dynamic key generation, structural enhancement, and dynamic S-box indexing. This work provides practical suggestions for designing secure chaos-based multimedia cryptosystems, offering valuable references for future research and promoting the development of efficient encryption systems. Yunan Wei, Donald Donglong Chen, Yupeng Li 0001, Ugur Erkan, Abdurrahim Toktas, Suo Gao, Yong Zhang 0018 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Scale Margin Loss for Object Detection
Yuxuan Cheng, Yanjun Zhang 0002, Leo Yu Zhang, Donald Donglong Chen, Yuming Fang 0001 |
KSEM (2) | 4 |
| 2025 | Using 3D-LMM-Based Encryption to Secure Digital Images With 3-D S-Box and Fibonacci Q-MatrixabstractThe rapid development of communication technology has significantly improved information transmission and increased capacity. In the context of the Internet of Things (IoT), where massive visual data transmission faces challenges of real-time processing and security threats. To address this problem, this paper proposes a novel encryption algorithm based on a 3D S-box combined with Fractal-Sort-Matrix (FSM) and Fibonacci Q-Matrix (3DSFF). To address the shortcomings of traditional S-box encryption, this paper integrates the 3D S-box with the FSM, thus enhancing the uncertainty of permutation and transformation. Additionally, to overcome the limitation of using a single value inciting Fibonacci Q-Matrix (FQM) applications, this paper combines FQM with chaotic sequences to strengthen its resistance against exhaustive attacks. Through comprehensive experimental tests on the algorithm’s outcomes, it demonstrates notable improvements over previous algorithms, achieving an average information entropy of 7.9993 and a correlation coefficient close to 0.01 after encryption. These tests indicate that the scheme can withstand common attacks, making it a sufficiently secure solution for the confidentiality of private images. Moreover, these advances provide a secure and efficient visual data protection framework for IoT applications involving anti-theft surveillance, data acquisition, and communication transmission. Yunlong Liao, Qiutong Li, Guoheng Huang, Donald Donglong Chen, Xiaochen Yuan |
IEEE Internet Things J. | 6 |
| 2025 | High-Radix/Mixed-Radix NTT Multiplication Algorithm/Architecture Co-Design Over Fermat ModulusabstractPolynomial multiplication using Number Theoretic Transform (NTT) is crucial in lattice-based post-quantum cryptography (PQC) and fully homomorphic encryption (FHE), with modulusqsignificantly affecting performance. Fermat moduli of the form$2^{2^{n}} + 1$, such as 65537, offer efficiency gains due to simplified modular reduction and powers-of-2 twiddle factors in NTT. While Fermat moduli have been directly applied or explored for incorporation into existing schemes, Fermat NTT-based polynomial multiplication designs remain underexplored in fully exploiting the benefits of Fermat moduli. This work presents a high-radix/mixed-radix NTT architecture tailored for Fermat moduli, which improves the utilization of the powers-of-2 twiddle factors in large transform sizes. In most cases, our design achieves a 30%–85% reduction in DSP area-time product (ATP) and a 70%–100% reduction in BRAM ATP compared to state-of-the-art designs with smaller or equivalent modulus, while maintaining competitive LUT and FF ATP, underscoring the potential of Fermat NTT-based polynomial multipliers in lattice-based cryptography. Yile Xing, Guangyan Li, Zewen Ye, Ryan W. L. Luk, Donald Donglong Chen, Hong Yan 0001, Ray C. C. Cheung |
IEEE Trans. Computers | 5 |
| 2025 | PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum ProcessorabstractPost-quantum cryptography (PQC) has rapidly evolved in response to the emergence of quantum computers, with the US National Institute of Standards and Technology (NIST) selecting four finalist algorithms for PQC standardization in 2022, including the Falcon digital signature scheme. Hawk is currently the only lattice-based candidate in NIST Round 2 additional signatures. Falcon and Hawk are based on the NTRU lattice, offering compact signatures, fast generation, and verification suitable for deployment on resource-constrained Internet-of-Things (IoT) devices. Despite the popularity of ML-DSA and ML-KEM, research on NTRU-based schemes has been limited due to their complex algorithms and operations. Falcon and Hawk's performance remains constrained by the lack of parallel execution in crucial operations like the Number Theoretic Transform (NTT) and Fast Fourier Transform (FFT), with data dependency being a significant bottleneck. This paper enhances NTRU-based schemes Falcon and Hawk through hardware/software co-design on a customized Single-Instruction-Multiple-Data (SIMD) processor, proposing new SIMD hardware units and instructions to expedite these schemes along with software optimizations to boost performance. Our NTT optimization includes a novel layer merging technique for SIMD architecture to reduce memory accesses, and the use of modular algorithms (Signed Montgomery and Improved Plantard) targets various modulus data widths to enhance performance. We explore applying layer merging to accelerate fixed-point FFT at the SIMD instruction level and devise a dual-issue parser to streamline assembly code organization to maximize dual-issue utilization. A System-on-chip (SoC) architecture is devised to improve the practical application of the processor in real-world scenarios. Evaluation on 28$nm$technology and field programmable gate array (FPGA) platform shows that our design and optimizations can increase the performance of Hawk signature generation and verification by over 7$\times$. Zewen Ye, Junhao Huang 0001, Tianshun Huang, Yudan Bai, Guangyan Li, Donald Donglong Chen, Ray C. C. Cheung, Kejie Huang |
IEEE Trans. Computers | 8 |
| 2024 | A4-Unet: Deformable Multi-Scale Attention Network for Brain Tumor SegmentationabstractBrain tumor segmentation models have aided diagnosis in recent years. However, they face MRI complexity and variability challenges, including irregular shapes and unclear boundaries, leading to noise, misclassification, and incomplete segmentation, thereby limiting accuracy. To address these issues, we adhere to an outstanding Convolutional Neural Networks (CNNs) design paradigm and propose a novel network named A4-Unet. In A4-Unet, Deformable Large Kernel Attention (DLKA) is incorporated in the encoder, allowing for improved capture of multi-scale tumors. Swin Spatial Pyramid Pooling (SSPP) with cross-channel attention is employed in a bottleneck further to study long-distance dependencies within images and channel relationships. To enhance accuracy, a Combined Attention Module (CAM) with Discrete Cosine Transform (DCT) orthogonality for channel weighting and convolutional element-wise multiplication is introduced for spatial weighting in the decoder. Attention gates (AG) are added in the skip connection to highlight the foreground while suppressing irrelevant background information. The proposed network is evaluated on three authoritative MRI brain tumor benchmarks and a proprietary dataset, and it achieves a 94.4% Dice score on the BraTS 2020 dataset, thereby establishing multiple new state-of-the-art benchmarks. The code is available here: https://github.com/WendyWAAAAANG/A4-Unet. Ruoxin Wang, Haiming Du, Yuxuan Cheng, Lingjie Yang, Xiaohui Duan, Yunfang Yu, Yu Zhou 0027, Donald Donglong Chen |
BIBM | 10 |
| 2024 | Unified Lossless-Throughput Architecture for AES and SM4 Encryption with Changeable KeysabstractNetwork devices targeting to implement data-intensive applications often require the outstanding performance of symmetric encryption, when dealing with multiple concurrent requests from multiple users. Despite the numerous works on high-performance implementation of AES and SM4, hybrid architectures with lossless throughput when the key changes have not been proposed. In this paper, we propose a unified fully-pipelined architecture of AES and SM4 targeting high-performance Galois/Counter Mode application scenarios. The architecture is able to maintain the consistent throughput of input and output datastreams with changeable keys. Compared with state-of-the-art works implemented with the TSMC 65nm process, our design can reduce the area by 26.83% by using a shared composite S-box. With the one-hot S-box, our design can reduce power consumption by 30.93% and increase throughput by 34.21%. Zhishuo Huang, Haosong Zhao, Donald Donglong Chen, Shuyan Zhu, Yinjin Fu, Nong Xiao 0001, Yao Liu 0006 |
ISCAS | 4 |
| 2024 | Multi-way High-Throughput Implementation of Kyber
Jipeng Zhang 0001, Junhao Huang 0001, Donald Donglong Chen, Lu Zhou 0002 |
ISC (2) | 4 |
| 2024 | ENG25519: Faster TLS 1.3 handshake using optimized X25519 and Ed25519
Jipeng Zhang 0001, Junhao Huang 0001, Lirui Zhao, Donald Donglong Chen, Çetin Kaya Koç |
USENIX Security Symposium | 4 |
| 2024 | Yet Another Improvement of Plantard Arithmetic for Faster Kyber on Low-End 32-bit IoT DevicesabstractIn 2022, the National Institute of Standards and Technology (NIST) made an announcement regarding the standardization of Post-Quantum Cryptography (PQC) candidates. Out of all the Key Encapsulation Mechanism (KEM) schemes, the CRYSTAL-Kyber emerged as the sole winner. This paper presents another improved version of Plantard arithmetic that could speed up Kyber implementations on two low-end 32-bit IoT platforms (ARM Cortex-M3 and RISC-V) without SIMD extensions. Specifically, we further enlarge the input range of the Plantard arithmetic without modifying its computation steps. After tailoring the Plantard arithmetic for Kyber’s modulus, we show that the input range of the Plantard multiplication by a constant is at least 2.14× larger than the original design in TCHES2022. Then, two optimization techniques for efficient Plantard arithmetic on Cortex-M3 and RISC-V are presented.We show that the Plantard arithmetic supersedes both Montgomery and Barrett arithmetic on low-end 32-bit platforms. With the enlarged input range and the efficient implementation of the Plantard arithmetic on these platforms, we propose various optimization strategies for NTT/INTT. We minimize or entirely eliminate the modular reduction of coefficients in NTT/INTT by taking advantage of the larger input range of the proposed Plantard arithmetic on low-end 32-bit platforms. Furthermore, we propose two memory optimization strategies that reduce 23.50%~28.31% stack usage for the speed-version Kyber implementation when compared to its counterpart on Cortex-M4. The proposed optimizations make the speed-version implementation more feasible on low-end IoT devices. Thanks to the aforementioned optimizations, our NTT/INTT implementation shows considerable speedups compared to the state-of-the-art work. Overall, we demonstrate the applicability of the speed-version Kyber implementation on memory-constrained IoT platforms and set new speed records for Kyber on these platforms. Junhao Huang 0001, Haosong Zhao, Jipeng Zhang 0001, Wangchen Dai, Lu Zhou 0002, Ray C. C. Cheung, Çetin Kaya Koç, Donald Donglong Chen |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2024 | ProgramGalois: A Programmable Generator of Radix-4 Discrete Galois Transformation Architecture for Lattice-Based CryptographyabstractLattice-based cryptography (LBC) has been established as a prominent research field, with particular attention on post-quantum cryptography (PQC) and fully homomorphic encryption (FHE). As the implementing bottleneck of PQC and FHE, number theoretic transform (NTT) has been extensively studied. However, current works struggled with scalability, hindering their adaptation to various parameters, such as bit width and polynomial length. In this article, we proposed a novel Discrete Galois Transformation (DGT) algorithm utilizing the radix-4 variant to achieve a higher level of parallelism to the existing NTT. Furthermore, to implement the efficient radix-4 DGT adapting more LBCs, we proposed a set of scalable building blocks, including a modified Barrett modular multiplier accepting arbitrary modulus with only one integer multiplier, a radix-4 DGT butterfly unit, and a stream permutation network. The proposed modules are implemented on the Xilinx Virtex-7 and U250 FPGA to evaluate resource utilization and performance. Lastly, a design space exploration framework is proposed to generate optimized radix-4 DGT hardware constrained by polynomial and platform parameters. The sensitivity analysis showcases the generated hardware’s performance and scalability. The implementation results on the Xilinx Virtex-7 and U250 FPGA show significant performance improvements over the state-of-the-art works, which reached at least 35%, 192%, and 68% area-time product improvements in terms of LUTs, BRAMs, and DSPs, respectively. Guangyan Li, Zewen Ye, Donald Donglong Chen, Wangchen Dai, Gaoyu Mao, Kejie Huang, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | GC-GAN: Photo Cartoonization Using Guided Cartoon Generative Adversarial Network
Huantong Hou, Jingjing Chen 0001, Donald Donglong Chen |
ICANN (5) | 4 |
| 2023 | High-performance and Configurable SW/HW Co-design of Post-quantum Signature CRYSTALS-DilithiumabstractCRYSTALS-Dilithium is a lattice-based post-quantum digital signature scheme that is resistant to attacks by quantum computers and has been selected to be standardized in the NIST post-quantum cryptography (PQC) standardization process. However, the speed performance and design flexibility of the Dilithium still need to be evaluated. This article presents a high-performance software/hardware co-design of CRYSTALS-Dilithium based on the NIST PQC round-3 parameters. High-speed pipelined hardware modules for NTT/INTT, point-wise multiplication/addition, and for SHAKE are included in the design to accelerate the time-consuming operations in Dilithium. All hardware modules are parameterized, thus allowing full support of runtime configuration to increase versatility. Moreover, the proposed software/hardware architecture and tight operating workflows reduce the data transmission overhead between the processor and other hardware modules. The hardware accelerator is implemented with a reconfigurable logic on FPGA and is integrated with the high-performance ARM Cortex-A9 processor in the Xilinx Zynq Architecture. We measure the performance of the software/hardware system for Dilithium in NIST security levels 2, 3, and 5. Compared to pure software implementations, we achieve 8.7–12.5 times speedup in Key generation, 6.3–7.3 times speedup in Sign, and 9.1–12.2 times speedup in Verify operations. Gaoyu Mao, Donald Donglong Chen, Guangyan Li, Wangchen Dai, Abdurrashid Ibrahim Sanka, Çetin Kaya Koç, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | DiGAN: Directional Generative Adversarial Network for Object TransfigurationabstractThe concept of cycle consistency in couple mapping has helped CycleGAN illustrate remarkable performance in the context of image-to-image translation. However, its limitations in object transfiguration have not been ideally solved yet. In order to alleviate previous problems of wrong transformation position, degeneration, and artifacts, this work presents a new approach called Directional Generative Adversarial Network (DiGAN) in the field of object transfiguration. The major contribution of this work is threefold. First, paired directional generators are designed for both intra-domain and inter-domain generations. Second, a segmentation network based on Mask R-CNN is introduced to build conditional inputs for both generators and discriminators. Third, a feature loss and a segmentation loss are added to optimize the model. Experimental results indicate that DiGAN surpasses CycleGAN and AttentionGAN by 17.2% and 60.9% higher on Inception Score, 15.5% and 2.05% lower on Fréchet Inception Distance, and 14.2% and 15.6% lower on VGG distance, respectively, in horse-to-zebra mapping. Yingfang Zhang, Peihao Zhong, Jingjing Chen 0001, Donald Donglong Chen |
ICMR | 5 |
| 2022 | High Throughput Hardware/Software Heterogeneous System for RRPN-Based Scene Text DetectionabstractRotation Region Proposal Networks (RRPN) are used to generate rotated proposals with the information of text angle for arbitrary oriented scene text detection (STD). However, the computational complexity of RRPN inference is relatively high compared with other methods, which makes it difficult for massive deployment. In this paper, the first full-stack FPGA-CPU heterogeneous system design of RRPN-based STD algorithm is proposed. A hardware/software partition method is presented to analyze and split the tasks to enhance the computation efficiency of hardware. The fast 2D Winograd algorithm and block floating point are utilized to reduce computation complexity while maintaining a relatively high precision. The implementation results show that the peak performance of MAC arrays in the proposed architecture reaches 655.4 GOPS and the energy efficiency achieves 64.9 GOPS/W. By fully exploiting the parallel and pipelined merits in the algorithms, the first hardware architectures for skew non-maximum suppression (S-NMS) layer and rotation region-of-interest (RRoI) polling layer are proposed. The throughput of the proposed hardware/software heterogeneous system achieves 40 times and 1.4 times improvements compared with CPU and GPU, respectively. Moreover, the comprehensive operating expense ratio of pure CPU, GPU, and the proposed system is 80.7:2.5:1, which indicates that it is suitable for massive deployment. Yao Xin, Donald Donglong Chen, Chongyang Zeng, Yi Wang 0004, Ray C. C. Cheung |
IEEE Trans. Computers | 2 |
| 2018 | FFT-Based McLaughlin's Montgomery Exponentiation without Conditional SelectionsabstractModular multiplication forms the basis of many cryptographic functions such as RSA, Diffie-Hellman key exchange, and ElGamal encryption. For large RSA moduli, combining the fast Fourier transform (FFT) with McLaughlin's Montgomery modular multiplication (MLM) has been validated to offer cost-effective implementation results. However, the conditional selections in McLaughlin's algorithm are considered to be inefficient and vulnerable to timing attacks, since extra long additions or subtractions may take place and the running time of MLM varies. In this work, we restrict the parameters of MLM by a set of new bounds and present a modified MLM algorithm involving no conditional selection. Compared to the original MLM algorithm, we inhibit extra operations caused by the conditional selections and accomplish constant running time for modular multiplications with different inputs. As a result, we improve both area-time efficiency and security against timing attacks. Based on the proposed algorithm, efficient FFT-based modular multiplication and exponentiation are derived. Exponentiation architectures with dual FFT-based multipliers are designed obtaining area-latency efficient solutions. The results show that our work offers a better efficiency compared to the state-of-the-art works from and above 2048-bit operand sizes. For single FFT-based modular multiplication, we have achieved constant running time and obtained area-latency efficiency improvements up to 24.3 percent for 1,024-bit and 35.5 percent for 4,096-bit operands, respectively. Wangchen Dai, Donald Donglong Chen, Ray C. C. Cheung, Çetin Kaya Koç |
IEEE Trans. Computers | 2 |
| 2017 | Area-Time Efficient Architecture of FFT-Based Montgomery MultiplicationabstractThe modular multiplication operation is the most time-consuming operation for number-theoretic cryptographic algorithms involving large integers, such as RSA and Diffie-Hellman. Implementations reveal that more than 75 percent of the time is spent in the modular multiplication function within the RSA for more than 1,024-bit moduli. There are fast multiplier architectures to minimize the delay and increase the throughput using parallelism and pipelining. However such designs are large in terms of area and low in efficiency. In this paper, we integrate the fast Fourier transform (FFT) method into the McLaughlin's framework, and present an improved FFT-based Montgomery modular multiplication (MMM) algorithm achieving high area-time efficiency. Compared to the previous FFT-based designs, we inhibit the zero-padding operation by computing the modular multiplication steps directly using cyclic and nega-cyclic convolutions. Thus, we reduce the convolution length by half. Furthermore, supported by the number-theoretic weighted transform, the FFT algorithm is used to provide fast convolution computation. We also introduce a general method for efficient parameter selection for the proposed algorithm. Architectures with single and double butterfly structures are designed obtaining low area-latency solutions, which we implemented on Xilinx Virtex-6 FPGAs. The results show that our work offers a better area-latency efficiency compared to the state-of-the-art FFT-based MMM architectures from and above 1,024-bit operand sizes. We have obtained area-latency efficiency improvements up to 50.9 percent for 1,024-bit, 41.9 percent for 2,048-bit, 37.8 percent for 4,096-bit and 103.2 percent for 7,680-bit operands. Furthermore, the operating latency is also outperformed with high clock frequency for length-64 transform and above. Wangchen Dai, Donald Donglong Chen, Ray C. C. Cheung, Çetin Kaya Koç |
IEEE Trans. Computers | 2 |
| 2016 | Parameter Space for the Architecture of FFT-Based Montgomery Modular MultiplicationabstractModular multiplication is the core operation in public-key cryptographic algorithms such as RSA and the Diffie-Hellman algorithm. The efficiency of the modular multiplier plays a crucial role in the performance of these cryptographic methods. In this paper, improvements to FFT-based Montgomery Modular Multiplication (FFTM3) using carry-save arithmetic and pre-computation techniques are presented. Moreover, pseudo-Fermat number transform is used to enrich the supported operand sizes for the FFTM3. The asymptotic complexity of our method is O(l log l log log l), which is the same as the Schonhage-Strassen multiplication algorithm (SSA). A systematic procedure to select suitable parameter set for the FFTM3is provided. Prototypes of the improved FFTM3multiplier with appropriate parameter sets are implemented on Xilinx Virtex-6 FPGA. Our method can perform 3,100-bit and 4,124-bit modular multiplications in 6.74 and 7.78 μs, respectively. It offers better computation latency and area-latency product compared to the state-of-the-art methods for operand size of 3,072-bit and above. Donald Donglong Chen, Gavin Xiaoxu Yao, Ray C. C. Cheung, Derek Chi-Wai Pao, Çetin Kaya Koç |
IEEE Trans. Computers | 1 |
| 2014 | Compact Ring-LWE Cryptoprocessor
Sujoy Sinha Roy, Frederik Vercauteren, Nele Mentens, Donald Donglong Chen, Ingrid Verbauwhede |
CHES | 4 |
| 2012 | Low complexity and hardware-friendly spectral modular multiplicationabstractThe Schönhage-Strassen Algorithm (SSA) is an asymptotically fast multiplication algorithm with the complexity of O(l log l log log l) where l is the operand size. It outperforms other multiplication algorithms when l is large enough. One possible usage of such long integer multiplication is for cryptography. Innovated from SSA, the Interleaved Spectral Montgomery Modular Multiplication (ISM3) algorithm is proposed to accelerate the modular multiplication. ISM3 algorithm primarily interleaves the Montgomery modular multiplication algorithm between time and spectral (frequency) domain. We show that the tasks in each step of the proposed algorithm have little data dependency, and hence, extremely suitable for hardware implementation. We present the parallel ISM3architecture and implement it on Xilinx Virtex-II and Virtex-6 FPGAs. Experimental results show that our 3838-bit ISM3 is faster than the previous Montgomery multiplier. Moreover, our design can complete a 7678-bit modular multiplication in 3398 cycles in 17.98 μs on a Virtex-6 device. Donald Donglong Chen, Gavin Xiaoxu Yao, Çetin Kaya Koç, Ray C. C. Cheung |
FPT | 1 |