Fangyu Zheng

dblp:153/3132 · DBLP profile ↗
← Back
46ranked-venue papers
2as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 22 · 1 first-author · 14 since 2021Systems, architecture and hardware · 12 · 1 first-author · 10 since 2021Computer networks · 8 · 6 since 2021Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Geng Yang 0001, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng, Meng Li 0004, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005
ISCA11
2026 GRASP: Accelerating Hash-Based PQC Performance on GPU Parallel Architecture
abstract
SPHINCS+, one of the Post-Quantum Cryptography Digital Signature Algorithms (PQC-DSA) selected by NIST in the third round, features very short public and private key lengths but faces significant performance challenges compared to other post-quantum cryptographic schemes, limiting its suitability for real-world applications. In scenarios involving a large number of concurrent signing or verification tasks, these performance bottlenecks become particularly critical. To address these challenges, we propose the GPU-based paRallel Accelerated SPHINCS+(GRASP), which leverages GPU technology to enhance the efficiency of SPHINCS+signing and verification processes. We propose an adaptable parallelization strategy for SPHINCS+, analyzing its signing and verification processes to identify critical sections for efficient parallel execution. Utilizing CUDA, we perform bottom-up optimizations, focusing on memory access patterns and hypertree computation, to enhance GPU resource utilization. These efforts, combined with kernel fusion technology, result in significant improvements in throughput and overall performance. Compared to previous works, our approach achieves the highest occupancy. Extensive experimentation demonstrates that our optimized CUDA implementation of SPHINCS+achieves superior performance. Specifically, our GRASP scheme delivers throughput improvements ranging from 1.09× to 3.45× compared to state-of-the-art GPU-based solutions and surpasses the NIST reference implementation by over three orders of magnitude, highlighting a significant performance advantage.
Yijing Ning, Jiankuo Dong, Jingqiang Lin 0001, Fangyu Zheng, Yu Fu 0007, Fu Xiao 0001
IEEE Trans. Computers4
2025 ML-Cube: Accelerating Module-Lattice-Based Cryptography using Machine Learning Accelerators with a Memory-Less Design
abstract
The rapid advancement of AI technologies has led to a dramatic surge in computational demands, driving significant breakthroughs in ML accelerators. The powerful performance of these accelerators has attracted the attention of cryptography researchers, and recent studies have begun to explore their use in accelerating cryptographic operations. However, treating these accelerators as black boxes leads to high latency, and strict concurrency requirements, which hinder their practical deployment. In this paper, we go beyond the black-box treatment of ML accelerators and introduce ML-Cube (ML3), a novel memory-less framework that leverages ML accelerators to implement module-lattice-based PQC, FIPS 203 ML-KEM, and FIPS 204 ML-DSA. The performance benefits of ML-Cube arise from our thorough analysis of ML accelerator internals. Rather than treating the accelerators as black boxes, we dissect their operating mechanisms and design tailored mathematical transformations for cryptographic acceleration. This enables memory-less (I)NTT and polynomial multiplication that minimizes external memory dependencies and reduces latency. We further address the high latency and excessive parallelism demands of traditional SIMT-based implementations by fully parallelizing both ML-KEM and ML-DSA schemes. Our experiments show that our Tensor Core-based (I)NTT achieves a 2.03x--3.56x speedup over a highly-optimized CUDA-core implementation. Moreover, our memory-less polynomial multiplication attains a 10x speedup, and the full ML-KEM reaches up to a 3.58x speedup with only less than one-tenth of the latency compared with SOTA approach (CHES '24). Additionally, our enhanced ML-DSA implementation offers a 30% to 55% throughput improvement over the previous SOTA methods (TDSC '24) under the server-oriented model. Importantly, by confining core computations within registers, our approach inherently mitigates memory disclosure and cache-based side-channel attacks, thereby enhancing overall security.
Fangyu Zheng, Zhuoyu Xie, Wenxu Tang, Guang Fan 0001, Yijing Ning, Yi Bian 0001, Jingqiang Lin 0001, Jiwu Jing
CCS2
2025 WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA Cores
abstract
The application of Fully Homomorphic Encryption (FHE) is rapidly gaining traction as a means to maintain data confidentiality while performing computations on encrypted data. Given the accessibility and computational power, GPUs hold promise for significantly accelerating FHE operations. However, existing GPU-based acceleration solutions face several formidable challenges, notably the extensive occurrence of pipeline stalls induced by memory access and suboptimal harnessing of GPU hardware. This paper presents WarpDrive, a comprehensive framework for GPU-based FHE acceleration. Through sophisticated computation decomposition and fine-grained memory access design, WarpDrive significantly reduces the number of instructions by $\mathbf{7 3 \%}$ and pipeline stalls by $\mathbf{8 6 \%}$ compared to the state-of-the-art solution. Additionally, WarpDrive features a framework that supports the concurrent utilization of CUDA Cores and Tensor Cores within the NTT operation, for the first time, achieving performance that surpasses that of any single type of processing unit. Furthermore, we fully exploit the intra-ciphertext parallelism to elevate both computation and memory utilization, achieving up to $2.12 \times$ improvements without the need for ciphertext batching. Experimental results demonstrate that our optimizations highly enhance the performance of homomorphic operations. On an NVIDIA A100 GPU, WarpDrive achieves a throughput of 1218 KOPS for NTT and 305 KOPS for homomorphic multiplication, outperforming the state-of-the-art GPU solution (TensorFHE) by factors of $13.4 \times$ and $3.5 \times$, respectively. For the specific FHE workload, even under a much smaller batch size, our approach achieves $2.8 \times$ the performance of TensorFHE.
Guang Fan 0001, Mingzhe Zhang 0005, Fangyu Zheng, Shengyu Fan, Xianglong Deng, Wenxu Tang, Liang Kong 0005, Shoumeng Yan
HPCA3
2025 AsyncGBP${}^{+}$+: Bridging SSL/TLS and Heterogeneous Computing Power With GPU-Based Providers
abstract
The rapid evolution of GPUs has emerged as a promising solution for accelerating the worldwide used SSL/TLS, which faces performance bottlenecks due to its underlying heavy cryptographic computations. Nevertheless, substantial structural adjustments from the parallel mode of GPUs to the serial mode of the SSL/TLS stack are imperative, potentially constraining the practical deployment of GPUs. In this paper, we propose AsyncGBP${}^{+}$, a three-level framework that facilitates the seamless conversion of cryptographic requests from synchronous to asynchronous mode. We conduct an in-depth analysis of the OpenSSL provider and cryptographic primitive features relevant to GPU implementations, aiming to fully exploit the potential of GPUs. Notably, AsyncGBP${}^{+}$supports three working settings (offline/online/hybrid), finely tailored for various public key cryptographic primitives, including traditional ones like X25519, Ed25519, ECDSA, and the quantum-safe CRYSTALS-Kyber. A comprehensive evaluation demonstrates that AsyncGBP${}^{+}$can efficiently achieve an improvement of up to 137.8$\times$compared to the default OpenSSL provider (for X25519, Ed25519, ECDSA) and 113.30$\times$compared to OpenSSL-compatibleliboqs(for CRYSTALS-Kyber) in a single-process setting. Furthermore, AsyncGBP${}^{+}$surpasses the current fastest commercial-off-the-shelf OpenSSL-compatible TLS accelerator with a 5.3$\times$to 7.0$\times$performance improvement.
Yi Bian 0001, Fangyu Zheng, Yuewu Wang, Lingguang Lei, Jiankuo Dong, Guang Fan 0001, Jiwu Jing
IEEE Trans. Computers2
2025 Revisiting Prediction-Based Min-Entropy Estimation: Toward Interpretability, Reliability, and Applicability
Dongchi Han, Tianyu Chen 0016, Shijie Jia 0001, Fangyu Zheng, Xianhui Lu
IEEE Trans. Inf. Forensics Secur.6
2025 HTM-PQC: Hardening Cryptography Keys Under the Trend of Post-Quantum Cryptography Migration on Industrial Internet
abstract
With the rapid expansion of Industry 4.0 technology, the proliferation of large-scale devices faces increasingly severe cyber threats, underscoring the critical importance of cryptographic technology for secure communication and authentication. However, cryptographic systems, as the bedrock of security, have faced a barrage of attacks in recent years, including potential threats from quantum computing and memory disclosure vulnerabilities. In this article, we focus on enhancing the security of two standard quantum-safe cryptographic algorithms, Dilithium and eXtended Merkle signature scheme (XMSS), by leveraging hardware transactional memory (HTM) to create a secure operational environment. Unlike traditional cryptography such as Rivest–Shamir–Adleman (RSA) and elliptic curve cryptography (ECC), Dilithium, and XMSS involve more and larger sensitive variables, rendering conventional solutions inadequate. By conducting a comprehensive sensitivity analysis of variables within the abovementioned algorithms, we confine sensitive operations to transactional execution regions and employ transaction-splitting technology for efficiency. Our prototype, utilizing Intel transactional synchronization extension (TSX), demonstrates robust protection against memory disclosure attacks with acceptable performance overheads. Notably, our security-enhanced Dilithium and XMSS software implementations, recommended by NIST, achieve an average throughput factor of 0.75 compared to the (unprotected) reference implementations.
Lingjia Meng, Yu Fu 0007, Fangyu Zheng, Ziqiang Ma, Jiankuo Dong, Jingqiang Lin 0001
IEEE Trans. Ind. Informatics3
2025 GIF-FHE: A Comprehensive Implementation and Evaluation of GPU-Accelerated FHE With Integer and Floating-Point Computing Power
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without revealing the plaintext, garnering significant interest from both academic and industrial communities. However, its broader adoption has been hindered by performance limitations. Consequently, researchers have turned to GPUs for efficient FHE implementation. Nevertheless, most have predominantly favored integer units due to their ease of use, overlooking the considerable computational potential of floating-point units in GPUs. Recognizing this untapped floating-point computational power, our paper introducesGIF-FHE, an extensive exploration and implementation of FHE, leveraging GPUs' integer and floating-point instructions for FHE acceleration. We develop a comprehensive suite of low-level and middle-level FHE primitives, offering multiple implementation variants with support for three word size configurations ($64/52/32$-bit). Particularly, we make innovative use of floating-point implementations, employing a novel methodology to efficiently leverage the floating-point unit's fused multiply-add (FMA) instructions. This represents the pioneering integration of floating-point units into FHE acceleration. To bridge our highly-optimized FHE primitives with practical applications, this paper also provides a high-level FHE implementation and interfaces that can be directly applied by upper-level applications such as neural network inference. Finally, we undertake a comprehensive experiment evaluation and comparison involving three types of arithmetic: FP64/INT64/INT32 with varying word size configurations and computation units. Notably, our fundamental function implementations consistently outperform counterparts on the same platform, achieving speedups ranging from$2.0\times$to$4.2\times$. In the context of CKKS FHE schemes, our homomorphic operation implementation surpasses the state-of-the-art GPU-based solution with a speedup of up to$3.8\times$, and exceeds the performance of the widely adopted CPU-based library, SEAL, with a remarkable speedup of over$300\times$.
Fangyu Zheng, Guang Fan 0001, Wenxu Tang, Yuan Zhao 0015, Jiankuo Dong, Jingqiang Lin 0001, Shoumeng Yan, Jiwu Jing
IEEE Trans. Parallel Distributed Syst.1
2024 CryptoPyt: Unraveling Python Cryptographic APIs Misuse with Precise Static Taint Analysis
abstract
Cryptographic APIs are essential for ensuring the security of software systems. However, many research studies have revealed that the misuse of cryptographic APIs is commonly widespread. Detecting such misuse in Python poses challenges due to its intricate features, including dynamic features and pass-by-object-reference. Existing tools lack the precision and accuracy to tackle these challenges, leading to both high false positives and false negatives. In this work, we propose a specific Python Cryptographic Abstract Syntax Tree (PCAST) to represent the structure of source code, which rewrites AST nodes to handle complex Python features. Based on PCAST, we design and implement CryptoPyt, a static code analysis tool that leverages precise taint analysis and 17 cryptographic misuse rules to automatically identify potential cryptographic APIs misuse in Python projects. We conduct an in-depth analysis of all the APIs within the popular 21 Python cryptographic libraries and design five kinds of taint detectors to perform intra-procedural and inter-function analysis on the APIs and arguments. To demonstrate the effectiveness of CryptoPyt, we conduct experiments with six state-of-the-art tools (i.e., Cryptolation, LICMA, Bandit, Dlint, Semgrep and CodeQL) on both the labeled benchmark PyCryptoBench and the real-world Python cryptographic projects datasets PCAMD. Our evaluations show that CryptoPyt achieves an F1 score of 0.80 on PyCryptoBench and a recall rate of 99.08% on PCAMD. Furthermore, we disclose the discovered critical issues to the developers and seven high-level CVE IDs have been assigned to these findings. Our tool contributes to enhancing the security of Python cryptographic software.
Xiangxin Guo, Shijie Jia 0001, Jingqiang Lin 0001, Fangyu Zheng, Guangzheng Li, Yueqiang Cheng, Kailiang Ji
ACSAC5
2024 DPad-HE: Towards Hardware-friendly Homomorphic Evaluation using 4-Directional Manipulation
abstract
Module Learning with Errors (MLWE) based approaches for Fully Homomorphic Encryption (FHE) have garnered attention due to their potential to enhance hardware-friendliness and implementation efficiency. However, despite these advantages, their overall performance still trails behind traditional schemes based on Ring Learning with Errors (RLWE). This indicates that while MLWE-based constructions hold promise, there remain significant challenges to overcome in bridging the performance gap with RLWE-based FHE schemes. By uncovering the reasons for the unsatisfactory performance of prior schemes and pinpointing the fundamental differences in the design of MLWE-based FHE compared to traditional approaches, the paper introduces DPad-HE with a novel design incorporating manipulation in the module rank dimension. The newly introduced operations, rank-up, and rank-down, effectively regulate the scale of gadget decomposition, reducing the computational workload of key-switching by several times. Taking CKKS as a case study, the evaluation showcases the comprehensive advantages of DPad-HE over the state-of-the-art MLWE-based scheme, resulting in a performance boost of 1.26× to 5.71×, a reduction in key size from 1/3 to 3/4, with enhanced noise control. To test the hardware-friendliness of the solution, DPad-HE is also implemented on GPU. Notably, DPad-HE demonstrates that, for the first time, the execution latency of MLWE-based schemes can achieve comparable performance with traditional RLWE ones, especially on the GPU platform where a speedup up to 1.41× is witnessed. Additionally, this paper provides a lightweight conversion method between RLWE and MLWE ciphertexts, allowing for flexible selection of RLWE and MLWE settings during a single complete evaluation process. This opens up new possibilities for both RLWE-based and MLWE-based FHEs.
Wenxu Tang, Fangyu Zheng, Guang Fan 0001, Jingqiang Lin 0001, Jiwu Jing
CCS2
2024 TLTracer: Dynamically Detecting Cache Side Channel Attacks with a Timing Loop Tracer
abstract
Recently, cache side-channel attacks have gained increasing attention due to the significant threat they pose to data security. As research advances, these attacks have become more covert and their impact has been widened. To mitigate the threat posed by cache side-channel attacks, numerous de-tection approaches have been proposed. However, they struggle to capture runtime features or depend heavily on hardware performance counters (HPCs), resulting in a significant number of false negatives or false positives. To address this issue, this paper proposes a broadly applicable runtime feature for identifying cache side-channel attack programs and introduces a dynamic binary analysis approach, TLTracer. TLTracer is runtime trace-based, independent of HPCs and capable of scanning and detecting whether a binary program is malicious before it is deployed in the real world. We implement a prototype of TLTracer and evaluate it with a set of malicious and benign programs. The results show that it can effectively detect the latest cache side-channel attacks without false positives, and offer increased resilience against adversarial evasion compared to other detection tools.
Lingjia Meng, Fangyu Zheng, Jingqiang Lin 0001, Shijie Jia 0001, Haoling Fan
ICC3
2024 TensorPolyMul: Accelerating Polynomial Multiplication in NTT-unfriendly Lattice-based Cryptography Using Tensor Cores
abstract
The urgent demand for computing power in Artificial intelligence (AI) technology has driven the rapid development of dedicated accelerators. Meanwhile, the threat posed by quantum computing to traditional public-key cryptography has prompted the emergence of post-quantum algorithms, such as lattice-based cryptography. However, performance issues with these algorithms have raised concerns within the industry about the transition to quantum-safe solutions. In this paper, we propose a novel universal framework for NTT-unfriendly lattice-based post-quantum algorithms, leveraging NVIDIA’s AI accelerator Tensor Core to address this challenge. By employing techniques such as polynomial matrixization and multi-precision representation, we effectively transform the primary workload (i.e., polynomial multiplication) into a series of small-coefficient matrix multiplications that can be directly accelerated by Tensor Cores. This approach effectively bridges the gap between typical Tensor Core workloads and the core workloads of lattice-based post-quantum cryptography. As a case study, we implemented a prototype called TensorPolyMul to provide an implementation of Saber, a quantum-safe Key Encapsulation Mechanism (KEM). The experiments showcase that TensorPolyMul surpasses the state-of-the-art Tensor Core-based work, achieving remarkable speed-ups of $1.53 \times 1.33 \times, 1.62 \times$, and $1.22 \times$ for Inner Product, MatrixVecMul, Encaps, and Decaps, respectively.
Yi Bian 0001, Fangyu Zheng, Jiwu Jing
ICPADS2
2024 TF-Timer: Mitigating Cache Side-Channel Attacks in Cloud through a Targeted Fuzzy Timer
abstract
Cache side-channel attacks pose a significant threat to the data security of multi-tenant public clouds. However, currently proposed defenses either lack transparency (requiring user involvement) or incur a significant performance penalty. This paper is motivated by our insightful observation for the behavior of cache side-channel attackers who employ rdtsc/rdtscp instructions for timing purposes. We have discerned a behavior pattern that enables comprehensive identification of potential attackers. Building upon this observation, we introduce TF -timer, which operates on the core principle of inspecting cache side-channel attacks using the pre-identified behavior pattern while obscuring the return values of rdtsc/rdtscp instructions. Our proposed technique preserves the properties of rdtsc/rdtscp, only blurring the attacker's timing to minimize the impact on other applications. We have implemented the prototype of TF-timer at the hypervisor layer. It is completely transparent to users and requires no hardware modifications. Our evaluation results demonstrate that TF -timer efficiently and precisely miti-gates cache side-channel attacks that exploit rdtsc/rdtscp for timing, with performance penalties within 1 %.
Shijie Jia 0001, Fangyu Zheng, Jingqiang Lin 0001, Lingjia Meng, Ziqiang Ma
WCNC3
2024 ZeroShield: Transparently Mitigating Code Page Sharing Attacks With Zero-Cost Stand-By
abstract
Numerous cache side-channel attack techniques enable attackers to execute a cross-VM cache side-channel attack through the sharing of code pages with the targeted victim. Nonetheless, most prior defense solutions fall short of efficiency and ease of deployment, thus restricting their practicality for real-world implementation. This paper introduces ZeroShield, an adaptive and transparent approach implemented at the hypervisor layer, designed to counteract the code page sharing attack, a subset of cache side-channel attacks, occurring within a single virtual machine (VM) or spanning across multiple VMs. By thoroughly scrutinizing the “by-products” resulting from a code page sharing attack, we meticulously track the attacker’s access to security-sensitive code pages. This is achieved through harnessing hardware virtualization features, such as the Intel extended page table, in conjunction with the CR3 register. Utilizing this information, ZeroShield continuously monitors security-sensitive code pages, adeptly navigating complex OS and hypervisor behaviors. The architecture of ZeroShield exhibits an attack-aware design, enabling it to deploy protection measures on demand. Consequently, the system theoretically experiences negligible overhead in the absence of attackers. Empirical evidence confirms the effectiveness of ZeroShield in thwarting code page sharing attacks. It achieves this without imposing any performance penalties in the absence of attackers, and with a minimal overhead of less than 3.8% when attackers are active. Significantly, ZeroShield boasts a cost-free standby state and necessitates no adjustments to upper applications, guest OS, or hardware configurations. This attribute positions ZeroShield as an optimal default solution in real-world cloud environments to effectively counter code page sharing attacks.
Fangyu Zheng, Jingqiang Lin 0001, Fangjie Jiang
IEEE Trans. Inf. Forensics Secur.2
2023 V-Curve25519: Efficient Implementation of Curve25519 on RISC-V Architecture
Qingguan Gao, Kaisheng Sun, Jiankuo Dong, Fangyu Zheng, Jingqiang Lin 0001, Yongjun Ren, Zhe Liu 0001
Inscrypt (2)4
2023 JWTKey: Automatic Cryptographic Vulnerability Detection in JWT Applications
Shijie Jia 0001, Jingqiang Lin 0001, Fangyu Zheng, Xiaozhuo Gu
ESORICS (3)4
2023 AsyncGBP: Unleashing the Potential of Heterogeneous Computing for SSL/TLS with GPU-based Provider
abstract
The proliferation of IoT and 5G technologies has led to an explosion of data traffic that data centers must handle while ensuring secure transmission via SSL/TLS. The high volume of cryptographic operations required imposes performance bottlenecks. The GPU-based cryptographic accelerator is one of the competitive solutions. However, significant structural differences with practical applications confine their capacities to specific domains, such as offline cryptanalysis, undermining their potential for real-world cryptographic acceleration.
Yi Bian 0001, Fangyu Zheng, Yuewu Wang, Lingguang Lei, Jiankuo Dong, Jiwu Jing
ICPP2
2023 Towards Faster Fully Homomorphic Encryption Implementation with Integer and Floating-point Computing Power of GPUs
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without knowledge of the plaintext message and currently has been the focus of both academia and industry. However, the performance issue hinders its large-scale application, highlighting the urgent requirements of high-performance FHE implementations.With noticing the tremendous potential of GPUs in the field of cryptographic acceleration, this paper comprehensively investigates how to convert the available computing resources residing in GPUs into FHE workhorses, and implement a full set of low-level and middle-level FHE primitives based on two arithmetic units (i.e., INT32 and FP64 units) with three types of data precision (i.e., INT32, INT64 and FP64). This paper gives a comprehensive evaluation and comparison based on each road-map. Our implementations of fundamental functions outperform the implementations on the same platform by 1.7× to 16.7×. Taking CKKS FHE schemes as a case study, our implementation of homomorphic multiplication achieves 3.2× speedup over the state-of-the-art GPU-based implementation, even considering the difference of platforms. The detailed evaluation and comparison of this paper would offer a vital reference for the follow-up work to choose appropriate underlying arithmetic units and important primitive optimizations in GPU-based FHE implementations.
Guang Fan 0001, Fangyu Zheng, Lipeng Wan 0002, Yuan Zhao 0015, Jiankuo Dong, Yuewu Wang, Jingqiang Lin 0001
IPDPS2
2023 Protecting Private Keys of Dilithium Using Hardware Transactional Memory
Lingjia Meng, Yu Fu 0007, Fangyu Zheng, Ziqiang Ma, Dingfeng Ye, Jingqiang Lin 0001
ISC3
2023 Hydamc: A Hybrid Detection Approach for Misuse of Cryptographic Algorithms in Closed-Source Software
abstract
Cryptographic algorithms are fundamental to secure software development, but security vulnerabilities can arise during implementation, usage, and when calling third-party libraries. As security standards continue to evolve, software updates have become an inevitable trend, and detecting cryptographic algorithm misuse is crucial to ensure compliance with these standards during the update process. However, closed-source software presents challenges in detecting cryptographic algorithm misuse. To enhance the security ecosystem of software, we designed a hybrid detection approach for detecting misuses in closed-source software related to weak cryptographic algorithms, short keys, insecure working modes, and insecure padding modes. Our hybrid detection tool uses both static and dynamic detection methods to collect log information through a logging mechanism in binary executable files. The collected data is cleaned using a data cleaning strategy and analyzed to extract key features, generating test reports to help developers and experts identify cryptographic algorithm security issues. We tested 24 software applications from app stores and found that 62.5% had weak algorithm implementations or usage, 83.3% supported short keys, and 50% supported insecure padding modes. Finally, we provided actionable recommendations to mitigate identified issues.
Haoling Fan, Fangyu Zheng, Jingqiang Lin 0001, Lingjia Meng, Shijie Jia 0001
TrustCom2
2023 EG-Four$\mathbb {Q}$: An Embedded GPU-Based Efficient ECC Cryptography Accelerator for Edge Computing
abstract
With the continuous development of Industry 4.0 technology, the embedded devices in Industrial Internet of Things (IIoT) are showing explosive growth, and large-scale cyber attacks or related security incidents continue to sound the alarm bell of information security. IIoT has strict requirements on computing performance and energy consumption, which poses severe challenges to cryptographic algorithms, especially public key cryptographic algorithms with high computational complexity. Embedded graphic processing unit (GPU) devices, always as edge computing nodes or AI accelerators, are widely deployed in IIoT applications. In this article, we propose an embedded GPU-based Four$\mathbb {Q}$(EG-Four$\mathbb {Q}$) elliptic curve public key cryptographic acceleration scheme. As far as we know, EG-Four$\mathbb {Q}$is the first work to completely implement Four$\mathbb {Q}$on the GPU platforms, including finite field operations, point arithmetic, and scalar multiplication. Relying only on 36-W power consumption, our scalar multiplication performance reaches 1717 kops/s with the latency of 2.38 ms. In terms of the energy-efficiency ratio, EG-Four$\mathbb {Q}$has significant advantages over other platforms such as advanced RISC machines (ARM) CPU, Intel CPU, field programmable gate array (FPGA), and desktop GPUs. The throughput of EG-Four$\mathbb {Q}$is 1.75 times that of the fastest elliptic curve cryptography implementation based on the same platform and even exceeds the performance of Intel top server CPU E5-2699v3 (18-core). Based on the embedded GPU Xavier, EG-Four$\mathbb {Q}$can act as a cryptographic edge computing module or even a cloud cryptographic accelerator, providing more efficient elliptic curve cryptographic services for IIoT.
Jiankuo Dong, Pinchang Zhang, Kaisheng Sun, Fu Xiao 0001, Fangyu Zheng, Jingqiang Lin 0001
IEEE Trans. Ind. Informatics5
2022 CryptoGo: Automatic Detection of Go Cryptographic API Misuses
abstract
Cryptographic algorithms act as essential ingredients of all secure systems. However, the expected security guarantee from cryptographic algorithms often falls short in practice due to various cryptographic application programming interfaces (API) misuses. While many research studies target cryptographic API misuses in the cases of Java, C/C++ and Python, similar issues within the Go domain are still uncovered.
Shijie Jia 0001, Fangyu Zheng, Jingqiang Lin 0001
ACSAC4
2022 A Novel High-Performance Implementation of CRYSTALS-Kyber with AI Accelerator
Lipeng Wan 0002, Fangyu Zheng, Guang Fan 0001, Rong Wei, Yuewu Wang, Jingqiang Lin 0001, Jiankuo Dong
ESORICS (3)2
2022 An Empirical Study on the Quality of Entropy Sources in Linux Random Number Generator
abstract
Random numbers are essential for communications security, as they are widely employed as secret keys and other critical parameters of cryptographic algorithms. The Linux random number generator (LRNG) is the most popular open-source software-based random number generator (RNG). The security of LRNG is influenced by the overall design, especially the quality of entropy sources. Therefore, it is necessary to assess and quantify the quality of the entropy sources which contribute the main randomness to RNGs. In this paper, we perform an empirical study on the quality of entropy sources in LRNG with Linux kernel 5.6, and provide the following two findings. We first analyze two important entropy sources: jiffies and cycles, and propose a method to predict jiffies by cycles with high accuracy. The results indicate that, the jiffies can be correctly predicted thus contain almost no entropy in the condition of knowing cycles. The other important finding is the failure of interrupt cycles during system boot. The lower bits of cycles caused by interrupts contain little entropy, which is contrary to our traditional cognition that lower bits have more entropy. We believe these findings are of great significance to improve the efficiency and security of the RNG design on software platforms.
Mingshu Du, Tianyu Chen 0016, Shijie Jia 0001, Fangyu Zheng
ICC6
2022 G-SM3: High-Performance Implementation of GPU-based SM3 Hash Function
abstract
Hash is one of the most important algorithms of cryptography, it is widely used in cryptographic primitives, such as digital signature, key exchange and so on. Further, hash cryptography is also the core operation of blockchain technology. With the explosive growth of the number of IoT devices and the rapid development of blockchain technology, the computing performance of hash has received widespread attention. The GPU high-performance computing platforms with a number of arithmetic cores are widely used in cryptographic optimization and acceleration. In this paper, we propose an efficient parallel accelerated framework of SM3 cryptography hash function based on GPU parallel computing devices, short for GPU-based SM3 (G-SM3). Our G-SM3 optimizes the implementation of the hash cryptographic algorithm from three aspects: parallelism, memory access and instructions. On the desktop GPU NVIDIA Titan V, the peak performance of G-SM3 reaches 23 GB/s, which is more than 7.5 times the performance of OpenSSL on a top-level server CPU (E5-2699V3) with 16 cores. On the embedded GPU which consumes less than 40 W, the SM3 throughput reaches 3.8 GB/s, which is even better than the performance of the serverlevel CPU. Based on the same GTX 1080, our performance is 1.12 times that of the fastest known GPU implementation, and the latency is reduced by more than 95%. Compared to other platforms, our G-SM3 has a huge advantage.
Jiankuo Dong, Pinchang Zhang, Fangyu Zheng, Fu Xiao 0001
ICPADS4
2022 TEGRAS: An Efficient Tegra Embedded GPU-Based RSA Acceleration Server
abstract
Industrial Internet of Things (IIoT) has strict requirements on the performance and security of devices. Public-key cryptography, as a kind of computing resource-consuming algorithm, is widely used in the digital signature, key exchange, and so on. The embedded graphics processing units (GPUs) are now rapidly achieving extraordinary computing power, such as NVIDIA Tegra K1/X1/X2/Xavier, which are also treated as edge computing devices. They are widely used in IIoT environments, such as intelligent manufacturing, smart cities, and vehicle-mounted systems. The performance advantages endow embedded GPUs with the possibility of accelerating cryptography that also requires high-density computing. This article implements an efficient Tegra-based embedded GPU RSA acceleration server-oriented IIoT, named TEGRAS. Various optimization methods are employed to promote efficiency, including multithreaded Montgomery multiplication and Chinese Remainder Theorem implementation on the resource-constricted embedded GPUs. With about 40–50 W of power consumption, TEGRAS can deliver 34 kops/s of RSA2048 signature generation and 1007 kops/s of RSA signature verification, which outperforms implementations in the desktop GPUs and embedded CPUs in the perspective of performance-to-power ratio. To evaluate TEGRAS in real-world scenarios, we additionally build a network stack to deliver digital signature services, which can provide more than 34 and 978 kops of signature generation and signature verification, respectively. In a word, based on the embedded GPU, we provide a high-throughput, low-latency, and ready-to-use RSA accelerator-oriented IIoT.
Jiankuo Dong, Guang Fan 0001, Fangyu Zheng, Tianyu Mao, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Internet Things J.3
2022 EC-ECC: Accelerating Elliptic Curve Cryptography for Edge Computing on Embedded GPU TX2
abstract
Driven by artificial intelligence and computer vision industries, Graphics Processing Units (GPUs) are now rapidly achieving extraordinary computing power. In particular, the NVIDIA Tegra K1/X1/X2 embedded GPU platforms, which are also treated as edge computing devices, are now widely used in embedded environments such as mobile phones, game consoles, and vehicle-mounted systems to support high-dimension display, auto-pilot, and so on. Meanwhile, with the rise of the Internet of Things (IoT), the demand for cryptographic operations for secure communications and authentications between edge computing nodes and IoT devices is also expanding. In this contribution, instead of the conventional implementations based on FPGA, ASIC, and ARM CPUs, we provide an alternative solution for cryptographic implementation on embedded GPU devices. Targeting the new cipher suite added in TLS 1.3, we implement Edwards25519/448 and Curve25519/448 on an edge computing platform, embedded GPU NVIDIA Tegra X2, where various performance optimizations are customized for the target platform, including a novel parallel method for the register-limited embedded GPUs. With about 15 W of power consumption, it can provide 210k/31k ops/s of Curve25519/448 scalar multiplication, 834k/123k ops/s of fixed-point Edwards25519/448 scalar multiplication, and 150k/22k ops/s of unknown-point one, which are respectively the primitives and main workloads of key agreement, signature generation, and verification of the TLS 1.3 protocol. Our implementations achieve 8 to 26 times speedup of OpenSSL running in the very powerful ARM CPU of the same platform and outperform the state-of-the-art implementations in FPGA by a wide margin with better power efficiency.
Jiankuo Dong, Fangyu Zheng, Jingqiang Lin 0001, Zhe Liu 0001, Fu Xiao 0001, Guang Fan 0001
ACM Trans. Embed. Comput. Syst.2
2021 Heterogeneous-PAKE: Bridging the Gap between PAKE Protocols and Their Real-World Deployment
abstract
Two entities, who only share a password and communicate over an insecure channel, authenticate each other and agree on a large session key for protecting their subsequent communication. This is called the password-authenticated key exchange (PAKE) protocol. PAKE protocol has been considered a suitable substitute for the prevailing hash-based authentication which is vulnerable to various attacks. However, vendors are discouraged by both its prohibitively computational overheads as well as integrating costs, leading to its limited use since being proposed.
Rong Wei, Fangyu Zheng, Jiankuo Dong, Guang Fan 0001, Lipeng Wan 0002, Jingqiang Lin 0001, Yuewu Wang
ACSAC2
2021 TESLAC: Accelerating Lattice-Based Cryptography with AI Accelerator
Lipeng Wan 0002, Fangyu Zheng, Jingqiang Lin 0001
SecureComm (1)2
2021 TX-RSA: A High Performance RSA Implementation Scheme on NVIDIA Tegra X2
Jiankuo Dong, Guang Fan 0001, Fangyu Zheng, Jingqiang Lin 0001, Fu Xiao 0001
WASA (2)3
2021 SECCEG: A Secure and Efficient Cryptographic Co-processor Based on Embedded GPU System
Guang Fan 0001, Fangyu Zheng, Jiankuo Dong, Jingqiang Lin 0001, Rong Wei, Lipeng Wan 0002
WASA (2)2
2021 DPF-ECC: A Framework for Efficient ECC With Double Precision Floating-Point Computing Power
abstract
Used ubiquitously in a huge amount of security protocols or applications, elliptic curve cryptography (ECC) is one of the most important cryptographic primitives, featuring efficiency and short key size compared with other public-key cryptosystems such as DSA and RSA. However, as a computation-intensive public-key cryptographic primitive, ECC arithmetic is still the bottleneck that restrains the overall performance of the end applications. In this paper, instead of the conventional and straightforward integer-based methods, we present a general framework to accelerate ECC schemes over prime field, called DPF-ECC, that deeply exploits double precision floating-point (DPF) computing power. The DPF-ECC framework finely manages each bit of the DPF numbers and minimizes the overhead brought by additional data format conversion, by making use of the DPF representation, the rounding operations, and fused multiply-add instruction supported by the IEEE 754 floating point standard. We also conduct two comprehensive case studies on Crandall primes and Solinas primes to demonstrate how the DPF-ECC framework is applied to the prevailing ECC schemes. To evaluate the proposed DPF-ECC framework in the real world, leveraging the floating-point computing power of GPUs, we implement Curve25519/448 and Edwards25519/448, the popular ECC schemes widely used in TLS 1.3, SSH, etc. The experimental result in Tesla P100 achieves a record-setting performance that outperforms the existing fastest integer work with 2x to 3x throughput. With dependency only on the very commonly supported IEEE 754 floating point standard, DPF-ECC framework can be a very competent and promising candidate for ECC implementation in most of general-purpose platforms.
Fangyu Zheng, Rong Wei, Jiankuo Dong, Niall Emmart, Jingqiang Lin 0001, Charles C. Weems
IEEE Trans. Inf. Forensics Secur.2
2020 SEGIVE: A Practical Framework of Secure GPU Execution in Virtualization Environment
abstract
With the advancement of processor technology, general-purpose GPUs have become popular parallel computing accelerators in the cloud. However, designed for graphics rendering and high-performance computing, GPUs are born without sound security mechanisms. Consequently, the GPU-based service in the cloud is vulnerable to attacks from the potentially compromised guest OS as large amounts of sensitive code and data are offloaded directly to the unprotected GPUs.In this paper, we propose SEGIVE, a practical framework of secure GPU execution in the virtualization environment, which protects offloaded device code and data from disclosure or tampering by malicious guest OSes through the full life cycle of security-critical GPU applications. First, SEGIVE secures all the traffic transferred to GPUs with Intel SGX technology, including the users' sensitive data and GPU binaries. Second, with various memory isolation mechanisms, SEGIVE enhances security in multi-user execution scenarios by sharing a GPU among multiple workloads, which avoids underutilization of device resources. Besides, SEGIVE requires no modifications to application source codes, the GPU architecture, or I/O interconnection to fulfill security principles, and thus almost all prevailing GPU-based applications can easily benefit from SEGIVE with little porting effort. We have implemented SEGIVE with KVM-QEMU on off-the-shelf NVIDIA GPUs and CPUs. Evaluation results show that with security-enhances, the performance of SEGIVE prototype is still competitive to the native execution on compute-intensive applications, especially for the public-key cryptography algorithm.
Fangyu Zheng, Jingqiang Lin 0001, Guang Fan 0001, Jiankuo Dong
IPCCC2
2020 DPF-ECC: Accelerating Elliptic Curve Cryptography with Floating-Point Computing Power of GPUs
abstract
Driven by artificial intelligence (AI) and computer vision industries, Graphics Processing Units (GPUs) are now rapidly achieving extraordinary computing power. In particular, the floating-point computing power, which is heavily relied on by graphics rendering and AI computation workload, is developing much faster in GPUs. Meanwhile, in many fields such as ecommerce and online finance, the demand for cryptographic operations for secure communications and authentication is also expanding.In this contribution, targeting the important cryptographic primitives widely used in TLS 1.3, etc., we implement Curve25519 and Edwards25519 with GPUs' floating-point computing power, where various performance optimization methods are customized for the target platform, including novel big-number representations combined with a new floating-point-based computing algorithm, efficient merged reduction strategies, and curve-level acceleration. This paper reports record-setting performance for the elliptic-curve method: on TITAN V, we respectively achieve 7.21 and 77.30 million operations per second of unknown and known point multiplication of Edwards25519, and 13.55 million operations per second of point multiplication of Curve25519. To the best of our knowledge, this contribution is the first to show that floating-point-based ECC implementations can outperform the integer-based ones by a huge margin. The experimental result in Tesla P100 achieves over double performance of the existing fastest integer work on the same platform, and the result in TITAN V sets a record for the throughput which is 4.43 times better than the second.
Fangyu Zheng, Niall Emmart, Jiankuo Dong, Jingqiang Lin 0001, Charles C. Weems
IPDPS2
2018 Faster Modular Exponentiation Using Double Precision Floating Point Arithmetic on the GPU
abstract
This paper presents a new approach to integer multiple precision (MP) modular exponentiation, using double-precision floating point (DPF) operations, that is suitable for GPU implementation. We show speedups ranging from 20 % to 34 % over the best prior GPU times for sizes corresponding to common RSA cryptographic operations (2048 to 4096 bits). Three techniques are described. First, by adding 2104to the high half of the product, and 252to the low half, we set the implicit leading 1 in the DPF mantissa so that the full 52 explicit bits are available for each half of the 104-bit products of samples. Second, the DPF values are cast bitwise to 64-bit integers for adding the column sums to get the MP result. Normally the cast would require masking off the exponents, but because they are constant, we can include them in the column sums and correct just once for their total. Third, by initializing the column sums with the appropriate negative value to compensate for the exponent sums, no corrective subtraction is needed. Our implementation on an NVIDIA GTX Titan Black GPU achieves between 132.5K and 161.9K modular exponentiations per second of size 1024 bits, with latencies ranging from 21.7 ms to 17.8 ms, making it practical for online RSA applications. Proportional results are shown for 1536 and 2048 bits. The implementation is so efficient that its maximum sustained performance is actually bounded by the thermal limit of the GPU.
Niall Emmart, Fangyu Zheng, Charles C. Weems
ARITH2
2018 A New Variant of the Barrett Algorithm Applied to Quotient Selection
abstract
Quotient Selection (QS) is a key step in the classic O(n2) multiple precision division algorithm. On processors with fast hardware division, it is a trivial problem, but on GPUs, division is quite slow. In this paper we investigate the effectiveness of Brent and Zimmermann's variant as well as our own novel variant of Barrett's algorithm. Our new approach is shown to be suitable for low radix (single precision) QS. Three highly optimized implementations, two of the Brent and Zimmerman variant and one based on our new approach, have been developed and we show that each is many times faster than using the division operation built in to the compiler. In addition, our variant is on average 22 % faster than the other two implementations. We also sketch proofs of correctness for all of the implementations and our new algorithm.
Niall Emmart, Fangyu Zheng, Charles C. Weems
ARITH2
2018 Utilizing GPU Virtualization to Protect the Private Keys of GPU Cryptographic Computation
Fangyu Zheng, Jingqiang Lin 0001, Jiankuo Dong
ICICS2
2018 sDPF-RSA: Utilizing Floating-point Computing Power of GPUs for Massive Digital Signature Computations
abstract
In financial, electronic and other security-sensitive industries, data centers require various protocols and algorithms to secure massive volumes of transactions. It is well known that digital signature is a computationally expensive task and a potential bottleneck that can restrict overall performance. In this paper, we make the following contributions. First, we propose a novel method called sDPF-RSA to accelerate the core algorithm of RSA, Montgomery multiplication, for Graphics Processing Units (GPUs). The sDPF approach takes advantage of the sign bit to increase the amount of information processed with each double precision floating point value and considerably improves performance. Second, we have comprehensively reviewed and tested the algorithms to ensure they all run in constant time. In particular we improve the standard carry resolution algorithm, introducing two constant time parallel techniques. We thus minimize the potential for timing attacks against GPU based RSA crypto-systems. Finally, we propose a full implementation of RSA, optimized for our GPU-accelerated computing platform to maximize its computing power. With protection against timing attacks, the throughputs of RSA-2048/3072/4096 on an NVIDIA GeForce GTX TITAN Black set a record of 52,747/15,179/6,435 (for signature generation) and 1,237,694/584,083/354,139 (for signature verification with public key 65,537) operations per second with modest latency, outperforming the contemporaneous CPU and many-core processor Xeon Phi by 3.9-11 times.
Jiankuo Dong, Fangyu Zheng, Niall Emmart, Jingqiang Lin 0001, Charles C. Weems
IPDPS2
2018 Building Your Private Cloud Storage on Public Cloud Service Using Embedded GPUs
Wangzhao Cheng, Fangyu Zheng, Wuqiong Pan, Jingqiang Lin 0001, Huorong Li, Bingyu Li 0003
SecureComm (1)2
2018 Secure and Efficient Outsourcing of Large-Scale Matrix Inverse Computation
Shiran Pan, Qiongxiao Wang, Fangyu Zheng, Jiankuo Dong
WASA3
2017 High-Performance Symmetric Cryptography Server with GPU Acceleration
Wangzhao Cheng, Fangyu Zheng, Wuqiong Pan, Jingqiang Lin 0001, Huorong Li, Bingyu Li 0003
ICICS2
2017 Utilizing the Double-Precision Floating-Point Computing Power of GPUs for RSA Acceleration
abstract
Asymmetric cryptographic algorithm (e.g., RSA and Elliptic Curve Cryptography) implementations on Graphics Processing Units (GPUs) have been researched for over a decade. The basic idea of most previous contributions is exploiting the highly parallel GPU architecture and porting the integer-based algorithms from general-purpose CPUs to GPUs, to offer high performance. However, the great potential cryptographic computing power of GPUs, especially by the more powerful floating-point instructions, has not been comprehensively investigated in fact. In this paper, we fully exploit the floating-point computing power of GPUs, by various designs, including the floating-point-based Montgomery multiplication/exponentiation algorithm and Chinese Remainder Theorem (CRT) implementation in GPU. And for practical usage of the proposed algorithm, a new method is performed to convert the input/output between octet strings and floating-point numbers, fully utilizing GPUs and further promoting the overall performance by about 5%. The performance of RSA-2048/3072/4096 decryption on NVIDIA GeForce GTX TITAN reaches 42,211/12,151/5,790 operations per second, respectively, which achieves 13 times the performance of the previous fastest floating-point-based implementation (published in Eurocrypt 2009). The RSA-4096 decryption precedes the existing fastest integer-based result by 23%.
Jiankuo Dong, Fangyu Zheng, Wuqiong Pan, Jingqiang Lin 0001, Jiwu Jing, Yuan Zhao 0015
Secur. Commun. Networks2
2017 An Efficient Elliptic Curve Cryptography Signature Server With GPU Acceleration
abstract
Over the Internet, digital signature has been an indispensable approach to securing e-commerce and other online transactions requiring authentication. Concerning the computing costs of signature generation and verification, it has become a more and more common practice for security practitioners to outsource such computations from heavily loaded application servers called tenants to dedicated proxies like signature servers in the enterprise private cloud. In this paper, we present our high-performance signature server called Guess. It implements the elliptic curve digital signature algorithm (ECDSA) with 256-b key size on a Linux-powered commodity computer, harnessing a desktop graphics processing unit as a featured cryptographic accelerator. We demonstrate our experience in maximizing the computing power of Guess and also its capability to deliver such power to the tenants, which includes down-to-earth customization and optimization considering various hardware and software factors. Our comprehensive implementation of ECDSA is tested against intensive network traffic. Field experiments show that Guess achieves Ts= 8.71 × 106operations per second (OPS) for signature generation or Tv= 9.29 × 105OPS for verification, which is significantly faster than existent prototypes and products. Guess is a universal server that readily supports various categories of elliptic curve cryptographic schemes, such as digital signature, key agreement, and encryption.
Wuqiong Pan, Fangyu Zheng, Wen Tao Zhu, Jiwu Jing
IEEE Trans. Inf. Forensics Secur.2
2016 PhiRSA: Exploiting the Computing Power of Vector Instructions on Intel Xeon Phi for RSA
Yuan Zhao 0015, Wuqiong Pan, Jingqiang Lin 0001, Peng Liu 0005, Fangyu Zheng
SAC6
2016 RegRSA: Using Registers as Buffers to Resist Memory Disclosure Attacks
Yuan Zhao 0015, Jingqiang Lin 0001, Wuqiong Pan, Fangyu Zheng, Ziqiang Ma
SEC5
2014 Exploiting the Floating-Point Computing Power of GPUs for RSA
Fangyu Zheng, Wuqiong Pan, Jingqiang Lin 0001, Jiwu Jing, Yuan Zhao 0015
ISC1