Jiankuo Dong

dblp:212/1277 · DBLP profile ↗
← Back
43ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0003-1693-3000ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 15 · 2 first-author · 13 since 2021Systems, architecture and hardware · 12 · 3 first-author · 10 since 2021Computer networks · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GRASP: Accelerating Hash-Based PQC Performance on GPU Parallel Architecture
abstract
SPHINCS+, one of the Post-Quantum Cryptography Digital Signature Algorithms (PQC-DSA) selected by NIST in the third round, features very short public and private key lengths but faces significant performance challenges compared to other post-quantum cryptographic schemes, limiting its suitability for real-world applications. In scenarios involving a large number of concurrent signing or verification tasks, these performance bottlenecks become particularly critical. To address these challenges, we propose the GPU-based paRallel Accelerated SPHINCS+(GRASP), which leverages GPU technology to enhance the efficiency of SPHINCS+signing and verification processes. We propose an adaptable parallelization strategy for SPHINCS+, analyzing its signing and verification processes to identify critical sections for efficient parallel execution. Utilizing CUDA, we perform bottom-up optimizations, focusing on memory access patterns and hypertree computation, to enhance GPU resource utilization. These efforts, combined with kernel fusion technology, result in significant improvements in throughput and overall performance. Compared to previous works, our approach achieves the highest occupancy. Extensive experimentation demonstrates that our optimized CUDA implementation of SPHINCS+achieves superior performance. Specifically, our GRASP scheme delivers throughput improvements ranging from 1.09× to 3.45× compared to state-of-the-art GPU-based solutions and surpasses the NIST reference implementation by over three orders of magnitude, highlighting a significant performance advantage.
Yijing Ning, Jiankuo Dong, Jingqiang Lin 0001, Fangyu Zheng, Yu Fu 0007, Fu Xiao 0001
IEEE Trans. Computers2
2026 Quantum-Resistant Data Sharing Scheme With Auditability for Internet of Vehicles
abstract
In the era of quantum computing, data sharing in the Internet of Vehicles (IoV) confronts the challenges of auditability, efficiency, and quantum security. However, existing research remains insufficient to meet the requirements of high mobility, resource constraints, and resilience against quantum attacks. In this paper, we propose a new quantum-secure auditable data sharing framework, in which we first present a quantum-resistant puncturable signature algorithm (QRPPRFS). Combining the low-noise LPN-based pseudorandom function with an optimized trapdoor generation mechanism, it achieves compact key sizes and millisecond-level signing; second, the blockchain and dual-commitment proof mechanism are integrated to ensure anonymity, transparent auditability and robustness. Finally, we rigorously demonstrate the correctness of our scheme, the EUF-CMA with puncturing of QRPPRFS, and the knowledge soundness and witness zero-knowledge of the dual-commitment proof system. Experimental evaluations show that, under the practical setting$n=256$and$q \approx 2^{23}$, the proposed scheme keeps both signing and verification latencies below 10 ms, and reduces the initial secret-key storage to only 0.22 MB. These results demonstrate that the proposed scheme achieves both enhanced security and high efficiency, outperforming existing schemes.
Lingyan Xue, Haiping Huang, Jiankuo Dong, Fu Xiao 0001
IEEE Trans. Dependable Secur. Comput.3
2026 X2O: Cross Parallel Optimization of the CROSS Post-Quantum Scheme on GPU
abstract
The CROSS Digital Signature Algorithm (DSA), currently a second-round candidate in the NIST standardization process for additional post-quantum digital signatures, offers compact public keys and strong security guarantees rooted in the code-based Restricted Syndrome Decoding Problem (R-SDP) and its variant R-SDP(G). Despite its strong theoretical foundation and practical significance, existing CPU-based implementations of CROSS exhibit evident performance limitations, while its potential for high-throughput acceleration on GPU architectures remains insufficiently investigated. In this work, we present X2O, the first systematically optimized GPU implementation framework for CROSS on NVIDIA GPUs. X2O introduces a novel cross-parallel architecture that integrates both horizontal and vertical parallelism to fully exploit the massive concurrency of modern GPU platforms. The framework incorporates a series of targeted optimizations, including fine-grained thread scheduling, optimized memory access patterns, hash function tuning, and GPU-efficient tree construction. Experimental results on a NVIDIA RTX 4090 demonstrate the efficiency of our design, achieving up to 1,082,904 signature generations and 1,589,595 verifications per second at NIST security level 1. Compared to the official AVX-optimized CPU implementation, our GPU-based approach achieves up to 120× speedup, establishing a new performance benchmark for CROSS and demonstrating the viability of high-throughput, post-quantum digital signatures on parallel computing platforms.
Yijing Ning, Jiankuo Dong, Jingqiang Lin 0001, Fu Xiao 0001
IEEE Trans. Inf. Forensics Secur.2
2026 HI-CKKS: Is High-Throughput Neglected? Reimagining CKKS Efficiency With Parallelism
abstract
The rapid advancement of the Industrial Internet of Things (IIoT) has positioned data privacy protection as a critical challenge in smart manufacturing. The Cheon–Kim–Kim–Song (CKKS) homomorphic encryption scheme, with its floating-point approximation and single instruction multiple data support capabilities, has emerged as a vital technology for IIoT privacy preservation. However, the high computational complexity of homomorphic multiplication significantly constrains its application performance in real-world industrial scenarios. While existing optimizations primarily focus on latency reduction, research on high-throughput parallel optimization remains notably insufficient. To address this, we propose HI-CKKS, a GPU-based high-throughput CKKS homomorphic multiplication optimization scheme for IIoT, achieving performance breakthroughs through multilevel technical innovations. By designing a batch-processing asynchronous execution architecture, we resolve host-server interaction delays in edge-cloud collaboration. A hierarchical hybrid optimization strategy combining instruction set acceleration, kernel fusion, and dynamic memory management significantly enhances (I)NTT operation efficiency. In addition, we construct a multidimensional parallel homomorphic multiplication model tailored for IIoT high-concurrency task characteristics. Experimental results show that HI-CKKS improves the throughput of NTT, INTT, and homomorphic multiplication by$175.08\times$,$191.27\times$, and$679.57\times$over CPU baselines. When compared to state-of-the-art GPU implementations on similar platforms, HI-CKKS attains$1.54\times$–$14.70\times$speedup for (I)NTT and$55.8\times$–$266.87\times$speedup for homomorphic multiplication. Our work provides an efficient solution for secure computation in IIoT and offers a new approach to performance optimization of homomorphic encryption in edge-cloud collaborative scenarios.
Fuyuan Chen, Jiankuo Dong, Zhenjiang Dong, Wangchen Dai
IEEE Trans. Ind. Informatics2
2026 RIGHT: GPU-Optimized Parallel PQC HAETAE for High-Throughput Cryptographic Acceleration
abstract
The rapid development of quantum computing poses a significant threat to the security of the Industrial Internet of Things (IIoT), rendering traditional cryptographic systems inadequate for ensuring long-term security. As a fundamental technology for establishing trust in IIoT networks, digital signatures are essential for secure device authentication, data integrity verification, and the protection of communication channels. However, implementing efficient digital signature schemes in resource-constrained embedded devices presents significant challenges, particularly when faced with the demands of postquantum cryptography (PQC). We propose a scheme for resource-intensive GPU optimization for HAETAE throughput tuning (RIGHT). Specifically, we employ coarse-grained parallelism and kernel fusion to maximize the parallel processing capabilities of GPUs. Furthermore, we present a hierarchical data locality optimization method for NTT/INTT and FFT, and optimize the SHAKE256 using CUDA PTX instructions, further enhancing throughput to meet the high concurrency and computational performance requirements of IIoT applications. On the Jetson Xavier, the signing performance of HAETAE is 1.25 x, 1.09 x, and 1.31 x that of the Intel i7-10700 K CPU with AVX2, while verification performance is about 4 times faster, demonstrating low power consumption and high efficiency, making it suitable for edge computing. Specifically, on the NVIDIA RTX 4090, the signature throughput reaches 155 324 ops/s, and the verification throughput reaches 4 276 075 ops/s, showcasing significant parallel computing capabilities, which is ideal for large-scale digital signature verification tasks in IIoT.
Jiankuo Dong, Yuze Hou, Mengke Liu, Lunjie Li, Zhenjiang Dong
IEEE Trans. Ind. Informatics2
2025 READ: Resource efficient authentication scheme for digital twin edge networks
Kai Wang 0072, Jiankuo Dong, Yijie Xu, Xinyi Ji, Letian Sha, Fu Xiao 0001
Future Gener. Comput. Syst.2
2025 AB-DHD: An Attention Mechanism and Bi-Directional Gated Recurrent Unit Based Model for Dynamic Link Library Hijacking Vulnerability Discovery
Xiao Chen 0017, Letian Sha, Fu Xiao 0001, Jiaye Pan, Jiankuo Dong
J. Comput. Sci. Technol.5
2025 AsyncGBP${}^{+}$+: Bridging SSL/TLS and Heterogeneous Computing Power With GPU-Based Providers
abstract
The rapid evolution of GPUs has emerged as a promising solution for accelerating the worldwide used SSL/TLS, which faces performance bottlenecks due to its underlying heavy cryptographic computations. Nevertheless, substantial structural adjustments from the parallel mode of GPUs to the serial mode of the SSL/TLS stack are imperative, potentially constraining the practical deployment of GPUs. In this paper, we propose AsyncGBP${}^{+}$, a three-level framework that facilitates the seamless conversion of cryptographic requests from synchronous to asynchronous mode. We conduct an in-depth analysis of the OpenSSL provider and cryptographic primitive features relevant to GPU implementations, aiming to fully exploit the potential of GPUs. Notably, AsyncGBP${}^{+}$supports three working settings (offline/online/hybrid), finely tailored for various public key cryptographic primitives, including traditional ones like X25519, Ed25519, ECDSA, and the quantum-safe CRYSTALS-Kyber. A comprehensive evaluation demonstrates that AsyncGBP${}^{+}$can efficiently achieve an improvement of up to 137.8$\times$compared to the default OpenSSL provider (for X25519, Ed25519, ECDSA) and 113.30$\times$compared to OpenSSL-compatibleliboqs(for CRYSTALS-Kyber) in a single-process setting. Furthermore, AsyncGBP${}^{+}$surpasses the current fastest commercial-off-the-shelf OpenSSL-compatible TLS accelerator with a 5.3$\times$to 7.0$\times$performance improvement.
Yi Bian 0001, Fangyu Zheng, Yuewu Wang, Lingguang Lei, Jiankuo Dong, Guang Fan 0001, Jiwu Jing
IEEE Trans. Computers7
2025 ECO-CRYSTALS: Efficient Cryptography CRYSTALS on Standard RISC-V ISA
abstract
The field of post-quantum cryptography (PQC) is continuously evolving. Many researchers are exploring efficient PQC implementation on various platforms, including x86, ARM, FPGA, GPU, etc. In this paper, we present an Efficient CryptOgraphy CRYSTALS (ECO-CRYSTALS) implementation on standard 64-bit RISC-V Instruction Set Architecture (ISA). The target schemes are two winners of the National Institute of Standards and Technology (NIST) PQC competition: CRYSTALS-Kyber and CRYSTALS-Dilithium, where the two most time-consuming operations are Keccak and polynomial multiplication. Notably, this paper is the first highly-optimized assembly software implementation to deploy Kyber and Dilithium on the 64-bit RISC-V ISA. Firstly, we propose a better scheduling strategy for Keccak, which is specifically tailored for the 64-bit dual-issue RISC-V architecture. Our 24-round Keccak permutation (Keccak-$p$[1600,24]) achieves a 59.18% speed-up compared to the reference implementation. Secondly, we apply two modular arithmetic (Montgomery arithmetic and Plantard arithmetic) in the polynomial multiplication of Kyber and Dilithium to get a better lazy reduction. Then, we propose a flexible dual-instruction-issue scheme of Number Theoretic Transform (NTT). As for the matrix-vector multiplication, we introduce a row-to-column processing methodology to minimize the expensive memory access operations. Compared to the reference implementation, we obtain a speedup of 53.85%$\thicksim$85.57% for NTT, matrix-vector multiplication, and INTT in our ECO-CRYSTALS. Finally, the ECO-CRYSTALS implementation for key generation, encapsulation, and decapsulation in Kyber achieves 399k, 448k, and 479k cycles respectively, achieving speedups of 60.82%, 63.93%, and 65.56% compared to the NIST reference implementation. Similarly, the ECO-CRYSTALS implementation for key generation, sign, and verify in Dilithium reaches 1 364k, 3 191k, and 1 369k cycles, showcasing speedups of 54.84%, 64.98%, and 57.20%, respectively.
Xinyi Ji, Jiankuo Dong, Junhao Huang 0001, Zhijian Yuan, Wangchen Dai, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Trans. Computers2
2025 AWDP-Automated Windows Domain Penetration Framework With Deep Reinforcement Learning
abstract
Windows domain is regarded as a primary target for intranet penetration since a large amount of sensitive information is stored in such domain with Windows OS. However, penetration testing is a intricate and time-consuming task, which is usually dedicated to experienced experts. To alleviate and partially solve this problem, we hereby propose an automated Windows domain penetration testing framework (AWDP). Firstly, we establish the test scenario as a Markov Decision Process (MDP) and then design a simulator for the Windows Domain penetration testing with OpenAI's Gymnasium. Secondly, we implement our automated Windows Domain penetration approach with four sequential steps, collecting domain and host information, modeling with acquired data, discovering optimal attack path through Deep Q-Learning Network (DQN), and performing penetration testing actions. Finally, to validate the effects of the proposed method, we conduct tests in real deployed domains. Experimental results demonstrate that, the proposed models and algorithms in the AWDP framework exhibit robust performance. Moreover, the framework adapts to different environments with rational and efficient estimated attack paths, which eventually enables end-to-end automation of Windows Domain penetration testing.
Letian Sha, Xingpeng Huo, Fu Xiao 0001, Jiankuo Dong, Ziyue Su
IEEE Trans. Dependable Secur. Comput.4
2025 Enhanced Two-Way Privacy-Preserving PHY-Layer Authentication for UAV-Assisted MIMO Systems
abstract
Authentication is a crucial method for ensuring the security of the unmanned aerial vehicle (UAV)-assisted communication systems; however, it also raises concerns about the leakage of private data. To address this problem, this paper focuses on the problem of identity authentication with privacy-preserving consideration. We propose a two-way privacy-preserving physical layer (PHY-Layer) authentication framework in a UAV-assisted multiple-input multiple-output (MIMO) communication system, by exploiting the carrier frequency offset (CFO) characterizing UAV identity and CFO-based session keys constructed by elliptic curve cryptography (ECC) to encrypt data frames. In particular, we first employ a MOOSE algorithm to extract the hardware fingerprint feature related to UAV, and based on the extracted CFO feature parameters, we devise an ECC-based algorithm for a session key negotiation. To achieve identity validation for the network parties, we establish a two-way authentication framework based on the resulting CFO feature parameters, and apply the CFO-based session key to encrypt data frames for avoiding the leakage of private data. Moreover, we derive the closed-form analytical expressions for the probabilities of the detection and false alarm for the rigorous performance analysis. Finally, we provide extensive numerical results to validate the proposed theoretical model and demonstrate its effectiveness in both identity authentication and privacy preservation. In comparison with the prior scheme, the proposed framework has a better robustness and privacy.
Pinchang Zhang, Huangwenqing Shi, Jiankuo Dong, Ji He 0002, Xiaohong Jiang 0001, Fu Xiao 0001
IEEE Trans. Dependable Secur. Comput.3
2025 GOLF: Unleashing GPU-Driven Acceleration for FALCON Post-Quantum Cryptography
abstract
Quantum computers leverage qubits to solve certain computational problems significantly faster than classical computers. This capability poses a severe threat to traditional cryptographic algorithms, leading to the rise of post-quantum cryptography (PQC) designed to withstand quantum attacks. FALCON, a lattice-based signature algorithm, has been selected by the National Institute of Standards and Technology (NIST) as part of its post-quantum cryptography standardization process. However, due to the computational complexity of PQC, especially in cloud-based environments, throughput limitations during peak demand periods have become a bottleneck, particularly for FALCON. In this paper, we introduce GOLF (GPU-accelerated Optimization for Lattice-based FALCON), a novel GPU-based parallel acceleration framework for FALCON. GOLF includes algorithm porting to the GPU, compatibility modifications, multi-threaded parallelism with distinct data, single-thread optimization for single tasks, and specific enhancements to the Fast Fourier Transform (FFT) module within FALCON. Our approach achieves unprecedented performance in FALCON acceleration on GPUs, setting the highest throughput record in the history of FALCON digital signature generation and verification. On the NVIDIA RTX 4090, GOLF reaches a signature generation throughput of 420.25 kops/s and a signature verification throughput of 10,311.04 kops/s. These results represent a 58.05× / 73.14× improvement over the reference FALCON implementation and a 7.17× / 3.79× improvement compared to the fastest known GPU implementation to date. Additionally, since we have not modified the content of the algorithm, but only optimized its engineering implementation, the security of the algorithm has not changed, and the security of the original algorithm has been maintained. GOLF demonstrates that GPU acceleration is not only feasible for post-quantum cryptography but also crucial for addressing throughput bottlenecks in real-world applications.
Ruihao Dai, Jiankuo Dong, Mingrui Qiu, Zhenjiang Dong, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Symphony of Speeds: Harmonizing Classic McEliece Cryptography With GPU Innovation
abstract
The Classic McEliece key encapsulation mechanism (KEM), a candidate in the fourth-round post-quantum cryptography (PQC) standardization process by the National Institute of Standards and Technology (NIST), stands out for its conservative design and robust security guarantees. Its deployment is impeded by exceptionally large public and secret keys. Modern GPUs offer abundant parallelism and global memory, making them well suited to such key sizes and to high-throughput cryptographic workloads. However, there has not been a systematic implementation of Classic McEliece on GPU platforms. This paper presents the first high-performance implementation of Classic McEliece on NVIDIA GPUs. Firstly, we present the first GPU-based implementation of Classic McEliece, utilizing a “CPU-GPU” heterogeneous approach and a kernel fusion strategy. We significantly reduce global memory accesses, optimizing memory access patterns. This results in encapsulation and decapsulation performance of 28,628,195 ops/s and 3,051,701 ops/s, respectively, for McEliece348864. Secondly, core operations like Additive Fast Fourier Transforms (AFFT), and Transpose AFFT (TAFFT) are optimized. We introduce the concept of the (T)AFFT stepping chain and propose two universal schemes: Memory Access Stepping Strategy (MASS) and Layer-Fused Memory Access Stepping Strategy (LFMASS), which achieve a speedup of 30.56% and 38.37%, respectively, compared to the native GPU-based McEliece6960119 implementation. Thirdly, extensive experiments on the NVIDIA RTX4090 show significant performance gains, achieving up to 344× higher encapsulation and 125× higher decapsulation compared to the official CPU-based AVX implementation, decisively outperforming existing ARM Cortex-M4 and FPGA implementations.
Jiankuo Dong, Zhenjiang Dong, Dung Hoang Duong, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Trans. Inf. Forensics Secur.2
2025 RLP-ABE: Puncturable CP-ABE for Efficient User Revocation From Lattices in Cloud Storage
abstract
Cloud computing has become the predominant platform for data sharing due to its adaptability, cost-effectiveness, and ability to scale resources according to user demand. Ensuring secure and efficient data sharing has long been a central research focus, with attribute-based encryption (ABE) serving as a key cryptographic primitive. In real-world scenarios, user attributes often change, necessitating timely revocation of access rights. Common user revocation methods include direct and indirect revocation. Direct revocation is controlled by the data owner, who adds revocation information to a list and embeds it into ciphertext to revoke permissions. Indirect revocation is managed by an authorized authority or delegated third party, dynamically publishing revocation information and generating new keys and ciphertexts. Conventional direct and indirect revocation methods incur substantial communication and computation overheads, limiting their practical effectiveness, particularly in environments with frequent user access terminations. To address these challenges, we propose a novel puncturable ciphertext-policy ABE scheme based on lattice cryptography for user revocation, eliminating the need for key regeneration and revocation-list maintenance. The proposed approach effectively resists collusion, quantum, and chosen-plaintext attacks, and experimental evaluations demonstrate its advantages in storage consumption, communication cost, and computational overhead.
Huaqun Wang, Debiao He, Jiankuo Dong
IEEE Trans. Inf. Forensics Secur.4
2025 SFO-CID: Structural Feature Optimization Based Command Injection Vulnerability Discovery for Internet of Things
abstract
The rapid development of Industrial Internet of Things (IIoT) has raised wider concerns for security of IoT devices. Command injection (CI) vulnerabilities, prevalent in IoT devices, pose a severe risk for remote code execution. Traditional static detection methods suffer from high overhead and imprecision due to symbolic execution. Popular binary code similarity detection (BCSD) methods rely on Control Flow Graphs (CFGs) with redundant structures, resulting in low efficiency and accuracy. In addition, they struggle with cross-function issues. In this paper, we proposeSFO-CID, a novel structural feature optimization based command injection vulnerability discovery model for IoT devices. Through backward taint analysis, all CFGs of suspicious CI vulnerabilities within the target binary file are precisely obtained. A large amount of code unrelated to vulnerabilities is removed, and cross-function issues are covered, significantly optimizing the structural features of original CFGs. Neural networks generate embedding vectors for optimized CFGs, transforming CI vulnerability detection into a vector similarity comparison. A wealth of semantic information within the code context is automatically and efficiently captured, improving the accuracy of vulnerability detection. We collect real-world cross-platform IoT firmware as data sources for tests. Experiments show thatSFO-CIDoutperforms popular BCSD methods, such asGemini,IoTSeeker, andFIT, achieving the highest accuracy of 88.67$\%$in vulnerability detection. Compared to existing state-of-the-art static analysis methods, likeKARONTEandSaTC,SFO-CIDattains the highest precision at 88.43$\%$and F1-score at 86.29$\%$, and is less time-consuming. Until now, 8 high-risk unknown vulnerabilities have been discovered, including 5 cross-function cases, and corresponding CVE IDs were assigned.
Xiao Chen 0017, Letian Sha, Fu Xiao 0001, Jiankuo Dong
IEEE Trans. Ind. Informatics5
2025 DDCC: Synergizing Denoising Diffusion Probabilistic Models and Curriculum-Based Complexity Control for Insider Threat Detection
abstract
Insider threat detection aims to identify malicious activities by employees that may compromise the confidentiality, integrity, or availability of organizational data. Detecting insider threats poses unique challenges compared to external threats, as internal actors often possess authorized system access and familiarity with organizational systems, enabling them to execute attacks discreetly. This article presents a novel approach to insider threat detection, synergizing denoising diffusion probabilistic models (DDPM) with a curriculum-based complexity control strategy (DDCC). While DDPM has shown promise in anomaly detection, its application to insider threat detection remains relatively unexplored. The proposed approach leverages DDPM for context-aware behavioral modeling and anomaly scoring. The methodology encompasses three primary components: First, employee behavioral logs undergo fusion to aggregate context information across various time scales. Second, a curriculum learning module controls the training process, gradually exposing the model to increasingly complex samples. The training data organization progresses from simple to complex sequences, facilitating more effective feature learning. Third, the denoising diffusion probabilistic model reconstructs the fused employee behavioral sequences for anomaly detection. Experimental validation on the CMU CERT r5.2 and r6.2 datasets demonstrates robust performance in insider threat detection across multiple time granularities of aggregated employee logs. The results underscore the effectiveness of this approach in addressing the nuanced challenges associated with detecting internal threats, highlighting its potential for real-world deployment in organizational security frameworks.
Jiankuo Dong, Zhenjiang Dong, Fuyuan Chen
IEEE Trans. Ind. Informatics1
2025 HTM-PQC: Hardening Cryptography Keys Under the Trend of Post-Quantum Cryptography Migration on Industrial Internet
abstract
With the rapid expansion of Industry 4.0 technology, the proliferation of large-scale devices faces increasingly severe cyber threats, underscoring the critical importance of cryptographic technology for secure communication and authentication. However, cryptographic systems, as the bedrock of security, have faced a barrage of attacks in recent years, including potential threats from quantum computing and memory disclosure vulnerabilities. In this article, we focus on enhancing the security of two standard quantum-safe cryptographic algorithms, Dilithium and eXtended Merkle signature scheme (XMSS), by leveraging hardware transactional memory (HTM) to create a secure operational environment. Unlike traditional cryptography such as Rivest–Shamir–Adleman (RSA) and elliptic curve cryptography (ECC), Dilithium, and XMSS involve more and larger sensitive variables, rendering conventional solutions inadequate. By conducting a comprehensive sensitivity analysis of variables within the abovementioned algorithms, we confine sensitive operations to transactional execution regions and employ transaction-splitting technology for efficiency. Our prototype, utilizing Intel transactional synchronization extension (TSX), demonstrates robust protection against memory disclosure attacks with acceptable performance overheads. Notably, our security-enhanced Dilithium and XMSS software implementations, recommended by NIST, achieve an average throughput factor of 0.75 compared to the (unprotected) reference implementations.
Lingjia Meng, Yu Fu 0007, Fangyu Zheng, Ziqiang Ma, Jiankuo Dong, Jingqiang Lin 0001
IEEE Trans. Ind. Informatics6
2025 VRVul-Discovery: BiLSTM-based Vulnerability Discovery for Virtual Reality Devices in Metaverse
abstract
The rapid development of the metaverse has brought about numerous security challenges. Virtual Reality (VR) , as one of the core technologies, plays a crucial role in the metaverse. The security of VR devices directly impacts user authentication and privacy. Currently, no attention has been paid to the vulnerabilities and security risks of VR devices. This article employs a bi-layer BiLSTM neural network to conduct a root cause analysis for user authentication and scene interaction when users enter metaverse environment using VR devices. By establishing the mapping between vulnerable VR firmware file attributes and metaverse interaction scenarios, we implement a vulnerability discovery and verification prototype called VRVul-Discovery, based on the concept of vulnerability discovery. Experiment results demonstrate that VRVul-Discovery provides high-accuracy determinations of firmware vulnerability attributes and scenarios susceptible to hijacking. In the end, the prototype system discovers seven unknown vulnerabilities, all of which are authenticated.
Letian Sha, Xiao Chen 0017, Fu Xiao 0001, Zhangbo Long, Qianyu Fan, Jiankuo Dong
ACM Trans. Multim. Comput. Commun. Appl.7
2025 GIF-FHE: A Comprehensive Implementation and Evaluation of GPU-Accelerated FHE With Integer and Floating-Point Computing Power
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without revealing the plaintext, garnering significant interest from both academic and industrial communities. However, its broader adoption has been hindered by performance limitations. Consequently, researchers have turned to GPUs for efficient FHE implementation. Nevertheless, most have predominantly favored integer units due to their ease of use, overlooking the considerable computational potential of floating-point units in GPUs. Recognizing this untapped floating-point computational power, our paper introducesGIF-FHE, an extensive exploration and implementation of FHE, leveraging GPUs' integer and floating-point instructions for FHE acceleration. We develop a comprehensive suite of low-level and middle-level FHE primitives, offering multiple implementation variants with support for three word size configurations ($64/52/32$-bit). Particularly, we make innovative use of floating-point implementations, employing a novel methodology to efficiently leverage the floating-point unit's fused multiply-add (FMA) instructions. This represents the pioneering integration of floating-point units into FHE acceleration. To bridge our highly-optimized FHE primitives with practical applications, this paper also provides a high-level FHE implementation and interfaces that can be directly applied by upper-level applications such as neural network inference. Finally, we undertake a comprehensive experiment evaluation and comparison involving three types of arithmetic: FP64/INT64/INT32 with varying word size configurations and computation units. Notably, our fundamental function implementations consistently outperform counterparts on the same platform, achieving speedups ranging from$2.0\times$to$4.2\times$. In the context of CKKS FHE schemes, our homomorphic operation implementation surpasses the state-of-the-art GPU-based solution with a speedup of up to$3.8\times$, and exceeds the performance of the widely adopted CPU-based library, SEAL, with a remarkable speedup of over$300\times$.
Fangyu Zheng, Guang Fan 0001, Wenxu Tang, Yuan Zhao 0015, Jiankuo Dong, Jingqiang Lin 0001, Shoumeng Yan, Jiwu Jing
IEEE Trans. Parallel Distributed Syst.7
2024 Exploiting Carrier Frequency Offset and Phase Noise for Physical Layer Authentication in UAV-Aided Communication Systems
abstract
This paper exploits two intrinsic hardware-specific fingerprints in terms of carrier frequency offset (CFO) and phase noise (PHN) to propose a two-dimensional physical layer authentication (PLA) scheme in the unmanned aerial vehicle (UAV)-aided communication systems. By leveraging expectation conditional maximization (ECM), extended Kalman filtering (EKF) algorithms and binary hypothesis testing, we first extract the inherent hardware impairments of UAV-aided systems including CFO and PHN as PHY-layer fingerprints to establish an authentication framework. To accurately characterize authentication performance, we examine the hybrid Cramér-Rao lower bound (HCRLB) for individual estimators of CFO and PHN, and then theoretically derive the analytical expressions for the false alarm and detection probabilities by utilizing tools from statistical signal processing. Finally, extensive numerical results are provided to validate the correctness of the developed theoretical models and to illustrate the authentication performance of the proposed scheme under various system parameters.
Yulin Teng, Pinchang Zhang, Jiankuo Dong, Fu Xiao 0001
IEEE Trans. Commun.4
2024 ECO-BIKE: Bridging the Gap Between PQC BIKE and GPU Acceleration
abstract
Advancements in quantum computing pose a threat to public-key cryptosystems, leading to the development of post-quantum cryptography. NIST is standardizing candidate algorithms, with BIKE, a code-based key encapsulation mechanism, among those under consideration. Performance is crucial in NIST PQC standardization process, and researchers have introduced a range of optimization techniques for BIKE across various platforms. To the best of our knowledge, our Efficient CryptOgraphy BIKE (ECO-BIKE) represents the first attempt at optimizing the implementation of BIKE on GPU architecture. In this paper, we introduce a comprehensive construction of a 3-threading parallel architecture tailored for the BIKE cryptosystem. This architecture covers a range of computational tasks, addressing operations from low-level to high-level computations. These include a parallel dense polynomial multiplication scheme with a better memory access pattern and a better XOR calculation, which forms the basis for a comprehensive parallel execution framework for the entire BIKE algorithm. Targeted optimizations are implemented for specific modules (KEYGEN, ENCAPS, DECAPS), which collectively enhance the overall efficiency of the algorithm. Our ECO-BIKE exhibits exceptional throughput performance on the NVIDIA GeForce RTX 4090. In the 3-thread mode, the throughput of the KEYGEN, ENCAPS, and DECAPS modules reaches 24.033 kops/s, 277.789 kops/s, and 5.817 kops/s, respectively. Our proposed optimal parallel multiplication scheme achieves a significantly higher overall throughput of 481.302 kops/s. These results highlight the substantial computational advantages our approach provides for cryptographic workloads.
Jiankuo Dong, Yusheng Fu, Xusheng Qin, Zhenjiang Dong, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Trans. Inf. Forensics Secur.1
2024 AUTH: An Adversarial Autoencoder Based Unsupervised Insider Threat Detection Scheme for Multisource Logs
abstract
Deep learning has shown broad research prospects in addressing insider threats, a serious problem currently facing industrial information systems. Although deep learning is able to capture effective feature representations from complex multidimensional data, there are still issues such as strong stealth of insider threat behavior and the imbalance data that need to be solved. Therefore, we propose an adversarial Autoencoder based Unsupervised insider Threat detection scHeme (AUTH). Compared to other methods, AUTH fully considers the role of time feature and event feature in threat detection. In addition, in order to improve the performance of autoencoder models to detect covert threat behaviors, AUTH drives a temporal convolutional network and long short-term memory network-based Adversarial Autoencoder (TL-AAE). Generative Adversarial Theory is introduced to solve the problem of uncertainty in the latent feature of the encoder. Finally, with the sufficient experiments on public datasets, we demonstrate that the usefulness of adding time features and the proposed TL-AAE model to improve threat detection performance. Compared with the baseline, AUTH obtains the area under curve value of 0.932, which is 4.95% higher than the highest result obtained by the baseline. In addition, AUTH obtains the EER value of 0.146, which is 12.57% lower than the lowest result of the baseline.
Xingjian Zhu, Jiankuo Dong, Zhen-Guo Zhou, Zhenjiang Dong, Yanfei Sun, Moyu Wang
IEEE Trans. Ind. Informatics2
2024 HI-Kyber: A Novel High-Performance Implementation Scheme of Kyber Based on GPU
abstract
CRYSTALS-Kyber, as the only public key encryption (PKE) algorithm selected by the National Institute of Standards and Technology (NIST) in the third round, is considered one of the most promising post-quantum cryptography (PQC) schemes. Lattice-based cryptography uses complex discrete algorithm problems on lattices to build secure encryption and decryption systems to resist attacks from quantum computing. Performance is an important bottleneck affecting the promotion of post quantum cryptography. In this paper, we present a High-performance Implementation of Kyber (named HI-Kyber) on the NVIDIA GPUs, which can increase the key-exchange performance of Kyber to the million-level. Firstly, we propose a lattice-based PQC implementation architecture based on kernel fusion, which can avoid redundant global-memory access operations. Secondly, We optimize and implement the core operations of CRYSTALS-Kyber, including Number Theoretic Transform (NTT), inverse NTT (INTT), pointwise multiplication, etc. Especially for the calculation bottleneck NTT operation, three novel methods are proposed to explore extreme performance: the sliced layer merging (SLM), the sliced depth-first search (SDFS-NTT) and the entire depth-first search (EDFS-NTT), which achieve a speedup of 7.5%, 28.5%, and 41.6% compared to the native implementation. Thirdly, we conduct comprehensive performance experiments with different parallel dimensions based on the above optimization. Finally, our key exchange performance reaches 1,664 kops/s. Specifically, based on the same platform, our HI-Kyber is 3.52× that of the GPU implementation based on the same instruction set and 1.78× that of the state-of-the-art one based on AI-accelerated tensor core.
Xinyi Ji, Jiankuo Dong, Tonggui Deng, Pinchang Zhang, Jiafeng Hua, Fu Xiao 0001
IEEE Trans. Parallel Distributed Syst.2
2023 XPORAM: A Practical Multi-client ORAM Against Malicious Adversaries
Biao Gao, Shijie Jia 0001, Jiankuo Dong, Peixin Ren
Inscrypt (1)3
2023 V-Curve25519: Efficient Implementation of Curve25519 on RISC-V Architecture
Qingguan Gao, Kaisheng Sun, Jiankuo Dong, Fangyu Zheng, Jingqiang Lin 0001, Yongjun Ren, Zhe Liu 0001
Inscrypt (2)3
2023 AsyncGBP: Unleashing the Potential of Heterogeneous Computing for SSL/TLS with GPU-based Provider
abstract
The proliferation of IoT and 5G technologies has led to an explosion of data traffic that data centers must handle while ensuring secure transmission via SSL/TLS. The high volume of cryptographic operations required imposes performance bottlenecks. The GPU-based cryptographic accelerator is one of the competitive solutions. However, significant structural differences with practical applications confine their capacities to specific domains, such as offline cryptanalysis, undermining their potential for real-world cryptographic acceleration.
Yi Bian 0001, Fangyu Zheng, Yuewu Wang, Lingguang Lei, Jiankuo Dong, Jiwu Jing
ICPP6
2023 Towards Faster Fully Homomorphic Encryption Implementation with Integer and Floating-point Computing Power of GPUs
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without knowledge of the plaintext message and currently has been the focus of both academia and industry. However, the performance issue hinders its large-scale application, highlighting the urgent requirements of high-performance FHE implementations.With noticing the tremendous potential of GPUs in the field of cryptographic acceleration, this paper comprehensively investigates how to convert the available computing resources residing in GPUs into FHE workhorses, and implement a full set of low-level and middle-level FHE primitives based on two arithmetic units (i.e., INT32 and FP64 units) with three types of data precision (i.e., INT32, INT64 and FP64). This paper gives a comprehensive evaluation and comparison based on each road-map. Our implementations of fundamental functions outperform the implementations on the same platform by 1.7× to 16.7×. Taking CKKS FHE schemes as a case study, our implementation of homomorphic multiplication achieves 3.2× speedup over the state-of-the-art GPU-based implementation, even considering the difference of platforms. The detailed evaluation and comparison of this paper would offer a vital reference for the follow-up work to choose appropriate underlying arithmetic units and important primitive optimizations in GPU-based FHE implementations.
Guang Fan 0001, Fangyu Zheng, Lipeng Wan 0002, Yuan Zhao 0015, Jiankuo Dong, Yuewu Wang, Jingqiang Lin 0001
IPDPS6
2023 PHY-layer authentication exploiting CFO for smart healthcare systems with mmWave communication technology
Yulin Teng, Huangwenqing Shi, Pinchang Zhang, Jiankuo Dong, Fu Xiao 0001
Ad Hoc Networks4
2023 EG-Four$\mathbb {Q}$: An Embedded GPU-Based Efficient ECC Cryptography Accelerator for Edge Computing
abstract
With the continuous development of Industry 4.0 technology, the embedded devices in Industrial Internet of Things (IIoT) are showing explosive growth, and large-scale cyber attacks or related security incidents continue to sound the alarm bell of information security. IIoT has strict requirements on computing performance and energy consumption, which poses severe challenges to cryptographic algorithms, especially public key cryptographic algorithms with high computational complexity. Embedded graphic processing unit (GPU) devices, always as edge computing nodes or AI accelerators, are widely deployed in IIoT applications. In this article, we propose an embedded GPU-based Four$\mathbb {Q}$(EG-Four$\mathbb {Q}$) elliptic curve public key cryptographic acceleration scheme. As far as we know, EG-Four$\mathbb {Q}$is the first work to completely implement Four$\mathbb {Q}$on the GPU platforms, including finite field operations, point arithmetic, and scalar multiplication. Relying only on 36-W power consumption, our scalar multiplication performance reaches 1717 kops/s with the latency of 2.38 ms. In terms of the energy-efficiency ratio, EG-Four$\mathbb {Q}$has significant advantages over other platforms such as advanced RISC machines (ARM) CPU, Intel CPU, field programmable gate array (FPGA), and desktop GPUs. The throughput of EG-Four$\mathbb {Q}$is 1.75 times that of the fastest elliptic curve cryptography implementation based on the same platform and even exceeds the performance of Intel top server CPU E5-2699v3 (18-core). Based on the embedded GPU Xavier, EG-Four$\mathbb {Q}$can act as a cryptographic edge computing module or even a cloud cryptographic accelerator, providing more efficient elliptic curve cryptographic services for IIoT.
Jiankuo Dong, Pinchang Zhang, Kaisheng Sun, Fu Xiao 0001, Fangyu Zheng, Jingqiang Lin 0001
IEEE Trans. Ind. Informatics1
2022 A Novel High-Performance Implementation of CRYSTALS-Kyber with AI Accelerator
Lipeng Wan 0002, Fangyu Zheng, Guang Fan 0001, Rong Wei, Yuewu Wang, Jingqiang Lin 0001, Jiankuo Dong
ESORICS (3)8
2022 G-SM3: High-Performance Implementation of GPU-based SM3 Hash Function
abstract
Hash is one of the most important algorithms of cryptography, it is widely used in cryptographic primitives, such as digital signature, key exchange and so on. Further, hash cryptography is also the core operation of blockchain technology. With the explosive growth of the number of IoT devices and the rapid development of blockchain technology, the computing performance of hash has received widespread attention. The GPU high-performance computing platforms with a number of arithmetic cores are widely used in cryptographic optimization and acceleration. In this paper, we propose an efficient parallel accelerated framework of SM3 cryptography hash function based on GPU parallel computing devices, short for GPU-based SM3 (G-SM3). Our G-SM3 optimizes the implementation of the hash cryptographic algorithm from three aspects: parallelism, memory access and instructions. On the desktop GPU NVIDIA Titan V, the peak performance of G-SM3 reaches 23 GB/s, which is more than 7.5 times the performance of OpenSSL on a top-level server CPU (E5-2699V3) with 16 cores. On the embedded GPU which consumes less than 40 W, the SM3 throughput reaches 3.8 GB/s, which is even better than the performance of the serverlevel CPU. Based on the same GTX 1080, our performance is 1.12 times that of the fastest known GPU implementation, and the latency is reduced by more than 95%. Compared to other platforms, our G-SM3 has a huge advantage.
Jiankuo Dong, Pinchang Zhang, Fangyu Zheng, Fu Xiao 0001
ICPADS1
2022 TEGRAS: An Efficient Tegra Embedded GPU-Based RSA Acceleration Server
abstract
Industrial Internet of Things (IIoT) has strict requirements on the performance and security of devices. Public-key cryptography, as a kind of computing resource-consuming algorithm, is widely used in the digital signature, key exchange, and so on. The embedded graphics processing units (GPUs) are now rapidly achieving extraordinary computing power, such as NVIDIA Tegra K1/X1/X2/Xavier, which are also treated as edge computing devices. They are widely used in IIoT environments, such as intelligent manufacturing, smart cities, and vehicle-mounted systems. The performance advantages endow embedded GPUs with the possibility of accelerating cryptography that also requires high-density computing. This article implements an efficient Tegra-based embedded GPU RSA acceleration server-oriented IIoT, named TEGRAS. Various optimization methods are employed to promote efficiency, including multithreaded Montgomery multiplication and Chinese Remainder Theorem implementation on the resource-constricted embedded GPUs. With about 40–50 W of power consumption, TEGRAS can deliver 34 kops/s of RSA2048 signature generation and 1007 kops/s of RSA signature verification, which outperforms implementations in the desktop GPUs and embedded CPUs in the perspective of performance-to-power ratio. To evaluate TEGRAS in real-world scenarios, we additionally build a network stack to deliver digital signature services, which can provide more than 34 and 978 kops of signature generation and signature verification, respectively. In a word, based on the embedded GPU, we provide a high-throughput, low-latency, and ready-to-use RSA accelerator-oriented IIoT.
Jiankuo Dong, Guang Fan 0001, Fangyu Zheng, Tianyu Mao, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Internet Things J.1
2022 EC-ECC: Accelerating Elliptic Curve Cryptography for Edge Computing on Embedded GPU TX2
abstract
Driven by artificial intelligence and computer vision industries, Graphics Processing Units (GPUs) are now rapidly achieving extraordinary computing power. In particular, the NVIDIA Tegra K1/X1/X2 embedded GPU platforms, which are also treated as edge computing devices, are now widely used in embedded environments such as mobile phones, game consoles, and vehicle-mounted systems to support high-dimension display, auto-pilot, and so on. Meanwhile, with the rise of the Internet of Things (IoT), the demand for cryptographic operations for secure communications and authentications between edge computing nodes and IoT devices is also expanding. In this contribution, instead of the conventional implementations based on FPGA, ASIC, and ARM CPUs, we provide an alternative solution for cryptographic implementation on embedded GPU devices. Targeting the new cipher suite added in TLS 1.3, we implement Edwards25519/448 and Curve25519/448 on an edge computing platform, embedded GPU NVIDIA Tegra X2, where various performance optimizations are customized for the target platform, including a novel parallel method for the register-limited embedded GPUs. With about 15 W of power consumption, it can provide 210k/31k ops/s of Curve25519/448 scalar multiplication, 834k/123k ops/s of fixed-point Edwards25519/448 scalar multiplication, and 150k/22k ops/s of unknown-point one, which are respectively the primitives and main workloads of key agreement, signature generation, and verification of the TLS 1.3 protocol. Our implementations achieve 8 to 26 times speedup of OpenSSL running in the very powerful ARM CPU of the same platform and outperform the state-of-the-art implementations in FPGA by a wide margin with better power efficiency.
Jiankuo Dong, Fangyu Zheng, Jingqiang Lin 0001, Zhe Liu 0001, Fu Xiao 0001, Guang Fan 0001
ACM Trans. Embed. Comput. Syst.1
2021 Heterogeneous-PAKE: Bridging the Gap between PAKE Protocols and Their Real-World Deployment
abstract
Two entities, who only share a password and communicate over an insecure channel, authenticate each other and agree on a large session key for protecting their subsequent communication. This is called the password-authenticated key exchange (PAKE) protocol. PAKE protocol has been considered a suitable substitute for the prevailing hash-based authentication which is vulnerable to various attacks. However, vendors are discouraged by both its prohibitively computational overheads as well as integrating costs, leading to its limited use since being proposed.
Rong Wei, Fangyu Zheng, Jiankuo Dong, Guang Fan 0001, Lipeng Wan 0002, Jingqiang Lin 0001, Yuewu Wang
ACSAC4
2021 TX-RSA: A High Performance RSA Implementation Scheme on NVIDIA Tegra X2
Jiankuo Dong, Guang Fan 0001, Fangyu Zheng, Jingqiang Lin 0001, Fu Xiao 0001
WASA (2)1
2021 SECCEG: A Secure and Efficient Cryptographic Co-processor Based on Embedded GPU System
Guang Fan 0001, Fangyu Zheng, Jiankuo Dong, Jingqiang Lin 0001, Rong Wei, Lipeng Wan 0002
WASA (2)3
2021 DPF-ECC: A Framework for Efficient ECC With Double Precision Floating-Point Computing Power
abstract
Used ubiquitously in a huge amount of security protocols or applications, elliptic curve cryptography (ECC) is one of the most important cryptographic primitives, featuring efficiency and short key size compared with other public-key cryptosystems such as DSA and RSA. However, as a computation-intensive public-key cryptographic primitive, ECC arithmetic is still the bottleneck that restrains the overall performance of the end applications. In this paper, instead of the conventional and straightforward integer-based methods, we present a general framework to accelerate ECC schemes over prime field, called DPF-ECC, that deeply exploits double precision floating-point (DPF) computing power. The DPF-ECC framework finely manages each bit of the DPF numbers and minimizes the overhead brought by additional data format conversion, by making use of the DPF representation, the rounding operations, and fused multiply-add instruction supported by the IEEE 754 floating point standard. We also conduct two comprehensive case studies on Crandall primes and Solinas primes to demonstrate how the DPF-ECC framework is applied to the prevailing ECC schemes. To evaluate the proposed DPF-ECC framework in the real world, leveraging the floating-point computing power of GPUs, we implement Curve25519/448 and Edwards25519/448, the popular ECC schemes widely used in TLS 1.3, SSH, etc. The experimental result in Tesla P100 achieves a record-setting performance that outperforms the existing fastest integer work with 2x to 3x throughput. With dependency only on the very commonly supported IEEE 754 floating point standard, DPF-ECC framework can be a very competent and promising candidate for ECC implementation in most of general-purpose platforms.
Fangyu Zheng, Rong Wei, Jiankuo Dong, Niall Emmart, Jingqiang Lin 0001, Charles C. Weems
IEEE Trans. Inf. Forensics Secur.4
2020 SEGIVE: A Practical Framework of Secure GPU Execution in Virtualization Environment
abstract
With the advancement of processor technology, general-purpose GPUs have become popular parallel computing accelerators in the cloud. However, designed for graphics rendering and high-performance computing, GPUs are born without sound security mechanisms. Consequently, the GPU-based service in the cloud is vulnerable to attacks from the potentially compromised guest OS as large amounts of sensitive code and data are offloaded directly to the unprotected GPUs.In this paper, we propose SEGIVE, a practical framework of secure GPU execution in the virtualization environment, which protects offloaded device code and data from disclosure or tampering by malicious guest OSes through the full life cycle of security-critical GPU applications. First, SEGIVE secures all the traffic transferred to GPUs with Intel SGX technology, including the users' sensitive data and GPU binaries. Second, with various memory isolation mechanisms, SEGIVE enhances security in multi-user execution scenarios by sharing a GPU among multiple workloads, which avoids underutilization of device resources. Besides, SEGIVE requires no modifications to application source codes, the GPU architecture, or I/O interconnection to fulfill security principles, and thus almost all prevailing GPU-based applications can easily benefit from SEGIVE with little porting effort. We have implemented SEGIVE with KVM-QEMU on off-the-shelf NVIDIA GPUs and CPUs. Evaluation results show that with security-enhances, the performance of SEGIVE prototype is still competitive to the native execution on compute-intensive applications, especially for the public-key cryptography algorithm.
Fangyu Zheng, Jingqiang Lin 0001, Guang Fan 0001, Jiankuo Dong
IPCCC5
2020 DPF-ECC: Accelerating Elliptic Curve Cryptography with Floating-Point Computing Power of GPUs
abstract
Driven by artificial intelligence (AI) and computer vision industries, Graphics Processing Units (GPUs) are now rapidly achieving extraordinary computing power. In particular, the floating-point computing power, which is heavily relied on by graphics rendering and AI computation workload, is developing much faster in GPUs. Meanwhile, in many fields such as ecommerce and online finance, the demand for cryptographic operations for secure communications and authentication is also expanding.In this contribution, targeting the important cryptographic primitives widely used in TLS 1.3, etc., we implement Curve25519 and Edwards25519 with GPUs' floating-point computing power, where various performance optimization methods are customized for the target platform, including novel big-number representations combined with a new floating-point-based computing algorithm, efficient merged reduction strategies, and curve-level acceleration. This paper reports record-setting performance for the elliptic-curve method: on TITAN V, we respectively achieve 7.21 and 77.30 million operations per second of unknown and known point multiplication of Edwards25519, and 13.55 million operations per second of point multiplication of Curve25519. To the best of our knowledge, this contribution is the first to show that floating-point-based ECC implementations can outperform the integer-based ones by a huge margin. The experimental result in Tesla P100 achieves over double performance of the existing fastest integer work on the same platform, and the result in TITAN V sets a record for the throughput which is 4.43 times better than the second.
Fangyu Zheng, Niall Emmart, Jiankuo Dong, Jingqiang Lin 0001, Charles C. Weems
IPDPS4
2018 Utilizing GPU Virtualization to Protect the Private Keys of GPU Cryptographic Computation
Fangyu Zheng, Jingqiang Lin 0001, Jiankuo Dong
ICICS4
2018 sDPF-RSA: Utilizing Floating-point Computing Power of GPUs for Massive Digital Signature Computations
abstract
In financial, electronic and other security-sensitive industries, data centers require various protocols and algorithms to secure massive volumes of transactions. It is well known that digital signature is a computationally expensive task and a potential bottleneck that can restrict overall performance. In this paper, we make the following contributions. First, we propose a novel method called sDPF-RSA to accelerate the core algorithm of RSA, Montgomery multiplication, for Graphics Processing Units (GPUs). The sDPF approach takes advantage of the sign bit to increase the amount of information processed with each double precision floating point value and considerably improves performance. Second, we have comprehensively reviewed and tested the algorithms to ensure they all run in constant time. In particular we improve the standard carry resolution algorithm, introducing two constant time parallel techniques. We thus minimize the potential for timing attacks against GPU based RSA crypto-systems. Finally, we propose a full implementation of RSA, optimized for our GPU-accelerated computing platform to maximize its computing power. With protection against timing attacks, the throughputs of RSA-2048/3072/4096 on an NVIDIA GeForce GTX TITAN Black set a record of 52,747/15,179/6,435 (for signature generation) and 1,237,694/584,083/354,139 (for signature verification with public key 65,537) operations per second with modest latency, outperforming the contemporaneous CPU and many-core processor Xeon Phi by 3.9-11 times.
Jiankuo Dong, Fangyu Zheng, Niall Emmart, Jingqiang Lin 0001, Charles C. Weems
IPDPS1
2018 Secure and Efficient Outsourcing of Large-Scale Matrix Inverse Computation
Shiran Pan, Qiongxiao Wang, Fangyu Zheng, Jiankuo Dong
WASA4
2017 Utilizing the Double-Precision Floating-Point Computing Power of GPUs for RSA Acceleration
abstract
Asymmetric cryptographic algorithm (e.g., RSA and Elliptic Curve Cryptography) implementations on Graphics Processing Units (GPUs) have been researched for over a decade. The basic idea of most previous contributions is exploiting the highly parallel GPU architecture and porting the integer-based algorithms from general-purpose CPUs to GPUs, to offer high performance. However, the great potential cryptographic computing power of GPUs, especially by the more powerful floating-point instructions, has not been comprehensively investigated in fact. In this paper, we fully exploit the floating-point computing power of GPUs, by various designs, including the floating-point-based Montgomery multiplication/exponentiation algorithm and Chinese Remainder Theorem (CRT) implementation in GPU. And for practical usage of the proposed algorithm, a new method is performed to convert the input/output between octet strings and floating-point numbers, fully utilizing GPUs and further promoting the overall performance by about 5%. The performance of RSA-2048/3072/4096 decryption on NVIDIA GeForce GTX TITAN reaches 42,211/12,151/5,790 operations per second, respectively, which achieves 13 times the performance of the previous fastest floating-point-based implementation (published in Eurocrypt 2009). The RSA-4096 decryption precedes the existing fastest integer-based result by 23%.
Jiankuo Dong, Fangyu Zheng, Wuqiong Pan, Jingqiang Lin 0001, Jiwu Jing, Yuan Zhao 0015
Secur. Commun. Networks1