Guang Fan 0001

dblp:46/2041-1 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-5389-0823ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 13 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Security and privacy · 4 · 4 since 2021
YearPublicationVenuePosition
2026 A Framework for Developing and Optimizing Fully Homomorphic Encryption Programs on GPUs
abstract
In sensitive domains such as healthcare and finance, machine learning increasingly employs Fully Homomorphic Encryption (FHE) to secure both user data and models. Although FHE's intrinsic parallelism naturally aligns with GPU architectures, optimizing GPU kernels alone remains insufficient for efficient end-to-end FHE application development. The inherent complexity of FHE schemes and intricate GPU-specific details impede developers from focusing on high-level program logic. Additionally, FHE's high memory requirements, fine-grained memory operations, and redundant computations introduce further optimization challenges, resulting in inefficiencies even when GPU kernels are individually optimized. This paper introduces EasyFHE, a framework designed to simplify the development and optimization of GPU-accelerated FHE applications. Similar to PyTorch, EasyFHE provides high-level interfaces for defining computational logic while automatically handling low-level tasks, such as implementation selection and memory management. Furthermore, it incorporates an optimization framework that systematically addresses performance bottlenecks by applying tailored optimization passes during the lowering from high-level FHE programs to GPU kernels. Compared to state-of-the-art open-source GPU FHE libraries, EasyFHE uniquely supports FHE programs with memory requirements exceeding typical GPU capacities, achieving an average speedup of 2.88× with a peak of 4.39×.
Jianyu Zhao 0004, Xueyu Wu 0001, Guang Fan 0001, Mingzhe Zhang 0005, Shoumeng Yan, Lei Ju 0001, Zhuoran Ji
ASPLOS (2)3
2026 Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption Accelerator
abstract
Fully homomorphic encryption (FHE) enables computation on encrypted data without compromising privacy, positioning it as a promising solution for secure cloud computing. However, its substantial computational overhead impedes practical deployment, prompting the development of dedicated hardware accelerators. In practice, when deploying cryptographic algorithm optimizations on FHE accelerators, hardware constraints typically such as limited memory capacity, often lead to a disparity between theoretical algorithmic advantage and achievable hardware efficiency.
Liang Kong 0005, Xianglong Deng, Guang Fan 0001, Shengyu Fan, Yilan Zhu, Geng Yang 0001, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005
ASPLOS (2)3
2026 An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design Automation
abstract
Fully Homomorphic Encryption (FHE) has become a promising approach to protecting data privacy in emerging application scenarios. Unfortunately, FHE suffers from significant processing speed degradation compared to plaintext computation, with one of the primary bottlenecks being the time-consuming Number Theoretic Transform (NTT). Therefore, accelerating NTT to accommodate various FHE parameters is crucial to advancing FHE towards practical use. With highly reconfigurable and performant logical fabrics, Field Programmable Gate Arrays (FPGAs) have exhibited great potential in NTT acceleration. By decomposing large-point NTT with strong data dependency into independent and simple small-point NTTs, the emerging Ten-step NTT (TNTT) algorithms intuitively enable higher parallelism and thereby have the potential to explore better performance compared to traditional algorithms. However, our quantitative analysis reveals that TNTT exhibits significant performance degradation as parallelism increases due to additional varying-size transpositions and Hadamard products. This paper proposes AutoNest, an efficient and scalable hardware architecture, along with an accelerator auto-generation framework for TNTT. The proposed hardware architecture maximizes performance by 1) adopting a 2D block decomposition dataflow to address critical path delays in transpose logic, thereby improving clock frequency. 2) integrating algorithm-level costfree twiddle factor fusion to reduce the number of modular multiplications in Hadamard products, thereby allowing higher parallelism on chip. Moreover, we also deliver an accelerator generation framework conducting automated design space exploration to elaborate a performant TNTT architecture under the target FPGAs' resource budget for user-defined FHE parameters. Experimental results on the AMD-Xilinx U280 FPGA demonstrate that NTT accelerators generated by AutoNest achieve an average speedup of$2.31 \times$compared to prior designs.
Yilan Zhu, Geng Yang 0001, Xingyu Tian, Dilshan Kumarathunga, Liang Kong 0005, Xianglong Deng, Shengyu Fan, Guang Fan 0001, Guiming Shi, Bo Zhang 0098, Yisong Chang, Shoumeng Yan, Zhenman Fang, Mingzhe Zhang 0005
HPCA8
2026 HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Geng Yang 0001, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng, Meng Li 0004, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005
ISCA1
2026 MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access Optimization
Xianglong Deng, Guang Fan 0001, Shengyu Fan, Mingzhe Zhang 0005
ISCA4
2025 The Future of Fully Homomorphic Encryption System: From a Storage I/O Perspective
Erci Xu, Shengyu Fan, Xianglong Deng, Guiming Shi, Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Shoumeng Yan, Mingzhe Zhang 0005
APPT7
2025 ML-Cube: Accelerating Module-Lattice-Based Cryptography using Machine Learning Accelerators with a Memory-Less Design
abstract
The rapid advancement of AI technologies has led to a dramatic surge in computational demands, driving significant breakthroughs in ML accelerators. The powerful performance of these accelerators has attracted the attention of cryptography researchers, and recent studies have begun to explore their use in accelerating cryptographic operations. However, treating these accelerators as black boxes leads to high latency, and strict concurrency requirements, which hinder their practical deployment. In this paper, we go beyond the black-box treatment of ML accelerators and introduce ML-Cube (ML3), a novel memory-less framework that leverages ML accelerators to implement module-lattice-based PQC, FIPS 203 ML-KEM, and FIPS 204 ML-DSA. The performance benefits of ML-Cube arise from our thorough analysis of ML accelerator internals. Rather than treating the accelerators as black boxes, we dissect their operating mechanisms and design tailored mathematical transformations for cryptographic acceleration. This enables memory-less (I)NTT and polynomial multiplication that minimizes external memory dependencies and reduces latency. We further address the high latency and excessive parallelism demands of traditional SIMT-based implementations by fully parallelizing both ML-KEM and ML-DSA schemes. Our experiments show that our Tensor Core-based (I)NTT achieves a 2.03x--3.56x speedup over a highly-optimized CUDA-core implementation. Moreover, our memory-less polynomial multiplication attains a 10x speedup, and the full ML-KEM reaches up to a 3.58x speedup with only less than one-tenth of the latency compared with SOTA approach (CHES '24). Additionally, our enhanced ML-DSA implementation offers a 30% to 55% throughput improvement over the previous SOTA methods (TDSC '24) under the server-oriented model. Importantly, by confining core computations within registers, our approach inherently mitigates memory disclosure and cache-based side-channel attacks, thereby enhancing overall security.
Fangyu Zheng, Zhuoyu Xie, Wenxu Tang, Guang Fan 0001, Yijing Ning, Yi Bian 0001, Jingqiang Lin 0001, Jiwu Jing
CCS5
2025 WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA Cores
abstract
The application of Fully Homomorphic Encryption (FHE) is rapidly gaining traction as a means to maintain data confidentiality while performing computations on encrypted data. Given the accessibility and computational power, GPUs hold promise for significantly accelerating FHE operations. However, existing GPU-based acceleration solutions face several formidable challenges, notably the extensive occurrence of pipeline stalls induced by memory access and suboptimal harnessing of GPU hardware. This paper presents WarpDrive, a comprehensive framework for GPU-based FHE acceleration. Through sophisticated computation decomposition and fine-grained memory access design, WarpDrive significantly reduces the number of instructions by $\mathbf{7 3 \%}$ and pipeline stalls by $\mathbf{8 6 \%}$ compared to the state-of-the-art solution. Additionally, WarpDrive features a framework that supports the concurrent utilization of CUDA Cores and Tensor Cores within the NTT operation, for the first time, achieving performance that surpasses that of any single type of processing unit. Furthermore, we fully exploit the intra-ciphertext parallelism to elevate both computation and memory utilization, achieving up to $2.12 \times$ improvements without the need for ciphertext batching. Experimental results demonstrate that our optimizations highly enhance the performance of homomorphic operations. On an NVIDIA A100 GPU, WarpDrive achieves a throughput of 1218 KOPS for NTT and 305 KOPS for homomorphic multiplication, outperforming the state-of-the-art GPU solution (TensorFHE) by factors of $13.4 \times$ and $3.5 \times$, respectively. For the specific FHE workload, even under a much smaller batch size, our approach achieves $2.8 \times$ the performance of TensorFHE.
Guang Fan 0001, Mingzhe Zhang 0005, Fangyu Zheng, Shengyu Fan, Xianglong Deng, Wenxu Tang, Liang Kong 0005, Shoumeng Yan
HPCA1
2025 FAST: An FHE Accelerator for Scalable-parallelism with Tunable-bit
abstract
Fully Homomorphic Encryption (FHE) enables direct computation on encrypted data, providing substantial security advantages in cloud-based modern society.However, FHE suffers from significant computational overhead compared to plaintext computation, hindering its adoption in real-world applications.While many accelerators have been designed to address performance bottlenecks, most do not fully leverage cryptographic optimization technologies, leaving room for further performance enhancements.In this work, we propose FAST, an FHE accelerator incorporating recent cryptographic optimizations, including hoisting technology and the gadget decomposition key-switching method (named KLSS method).We analyze ciphertext level consumption throughout application execution and observe that workload requirements vary significantly with different ciphertext levels for both hybrid and KLSS key-switching methods.Additionally, we note the differing computational precision requirements for these key-switching methods.Based on these observations, we designed a versatile framework that supports multiple key-switching methods during a single application execution and integrates hoisting technology.
Shengyu Fan, Xianglong Deng, Liang Kong 0005, Guiming Shi, Guang Fan 0001, Dan Meng 0002, Rui Hou 0001, Mingzhe Zhang 0005
ISCA5
2025 HAWK: Fully Homomorphic Encryption Accelerator with Fixed-Word Key Decomposition Switching
Liang Kong 0005, Shengyu Fan, Xianglong Deng, Guang Fan 0001, Guiming Shi, Yilan Zhu, Geng Yang 0001, Shoumeng Yan, Mingzhe Zhang 0005
MICRO5
2025 AsyncGBP${}^{+}$+: Bridging SSL/TLS and Heterogeneous Computing Power With GPU-Based Providers
abstract
The rapid evolution of GPUs has emerged as a promising solution for accelerating the worldwide used SSL/TLS, which faces performance bottlenecks due to its underlying heavy cryptographic computations. Nevertheless, substantial structural adjustments from the parallel mode of GPUs to the serial mode of the SSL/TLS stack are imperative, potentially constraining the practical deployment of GPUs. In this paper, we propose AsyncGBP${}^{+}$, a three-level framework that facilitates the seamless conversion of cryptographic requests from synchronous to asynchronous mode. We conduct an in-depth analysis of the OpenSSL provider and cryptographic primitive features relevant to GPU implementations, aiming to fully exploit the potential of GPUs. Notably, AsyncGBP${}^{+}$supports three working settings (offline/online/hybrid), finely tailored for various public key cryptographic primitives, including traditional ones like X25519, Ed25519, ECDSA, and the quantum-safe CRYSTALS-Kyber. A comprehensive evaluation demonstrates that AsyncGBP${}^{+}$can efficiently achieve an improvement of up to 137.8$\times$compared to the default OpenSSL provider (for X25519, Ed25519, ECDSA) and 113.30$\times$compared to OpenSSL-compatibleliboqs(for CRYSTALS-Kyber) in a single-process setting. Furthermore, AsyncGBP${}^{+}$surpasses the current fastest commercial-off-the-shelf OpenSSL-compatible TLS accelerator with a 5.3$\times$to 7.0$\times$performance improvement.
Yi Bian 0001, Fangyu Zheng, Yuewu Wang, Lingguang Lei, Jiankuo Dong, Guang Fan 0001, Jiwu Jing
IEEE Trans. Computers8
2025 GIF-FHE: A Comprehensive Implementation and Evaluation of GPU-Accelerated FHE With Integer and Floating-Point Computing Power
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without revealing the plaintext, garnering significant interest from both academic and industrial communities. However, its broader adoption has been hindered by performance limitations. Consequently, researchers have turned to GPUs for efficient FHE implementation. Nevertheless, most have predominantly favored integer units due to their ease of use, overlooking the considerable computational potential of floating-point units in GPUs. Recognizing this untapped floating-point computational power, our paper introducesGIF-FHE, an extensive exploration and implementation of FHE, leveraging GPUs' integer and floating-point instructions for FHE acceleration. We develop a comprehensive suite of low-level and middle-level FHE primitives, offering multiple implementation variants with support for three word size configurations ($64/52/32$-bit). Particularly, we make innovative use of floating-point implementations, employing a novel methodology to efficiently leverage the floating-point unit's fused multiply-add (FMA) instructions. This represents the pioneering integration of floating-point units into FHE acceleration. To bridge our highly-optimized FHE primitives with practical applications, this paper also provides a high-level FHE implementation and interfaces that can be directly applied by upper-level applications such as neural network inference. Finally, we undertake a comprehensive experiment evaluation and comparison involving three types of arithmetic: FP64/INT64/INT32 with varying word size configurations and computation units. Notably, our fundamental function implementations consistently outperform counterparts on the same platform, achieving speedups ranging from$2.0\times$to$4.2\times$. In the context of CKKS FHE schemes, our homomorphic operation implementation surpasses the state-of-the-art GPU-based solution with a speedup of up to$3.8\times$, and exceeds the performance of the widely adopted CPU-based library, SEAL, with a remarkable speedup of over$300\times$.
Fangyu Zheng, Guang Fan 0001, Wenxu Tang, Yuan Zhao 0015, Jiankuo Dong, Jingqiang Lin 0001, Shoumeng Yan, Jiwu Jing
IEEE Trans. Parallel Distributed Syst.2
2024 DPad-HE: Towards Hardware-friendly Homomorphic Evaluation using 4-Directional Manipulation
abstract
Module Learning with Errors (MLWE) based approaches for Fully Homomorphic Encryption (FHE) have garnered attention due to their potential to enhance hardware-friendliness and implementation efficiency. However, despite these advantages, their overall performance still trails behind traditional schemes based on Ring Learning with Errors (RLWE). This indicates that while MLWE-based constructions hold promise, there remain significant challenges to overcome in bridging the performance gap with RLWE-based FHE schemes. By uncovering the reasons for the unsatisfactory performance of prior schemes and pinpointing the fundamental differences in the design of MLWE-based FHE compared to traditional approaches, the paper introduces DPad-HE with a novel design incorporating manipulation in the module rank dimension. The newly introduced operations, rank-up, and rank-down, effectively regulate the scale of gadget decomposition, reducing the computational workload of key-switching by several times. Taking CKKS as a case study, the evaluation showcases the comprehensive advantages of DPad-HE over the state-of-the-art MLWE-based scheme, resulting in a performance boost of 1.26× to 5.71×, a reduction in key size from 1/3 to 3/4, with enhanced noise control. To test the hardware-friendliness of the solution, DPad-HE is also implemented on GPU. Notably, DPad-HE demonstrates that, for the first time, the execution latency of MLWE-based schemes can achieve comparable performance with traditional RLWE ones, especially on the GPU platform where a speedup up to 1.41× is witnessed. Additionally, this paper provides a lightweight conversion method between RLWE and MLWE ciphertexts, allowing for flexible selection of RLWE and MLWE settings during a single complete evaluation process. This opens up new possibilities for both RLWE-based and MLWE-based FHEs.
Wenxu Tang, Fangyu Zheng, Guang Fan 0001, Jingqiang Lin 0001, Jiwu Jing
CCS3
2023 Towards Faster Fully Homomorphic Encryption Implementation with Integer and Floating-point Computing Power of GPUs
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without knowledge of the plaintext message and currently has been the focus of both academia and industry. However, the performance issue hinders its large-scale application, highlighting the urgent requirements of high-performance FHE implementations.With noticing the tremendous potential of GPUs in the field of cryptographic acceleration, this paper comprehensively investigates how to convert the available computing resources residing in GPUs into FHE workhorses, and implement a full set of low-level and middle-level FHE primitives based on two arithmetic units (i.e., INT32 and FP64 units) with three types of data precision (i.e., INT32, INT64 and FP64). This paper gives a comprehensive evaluation and comparison based on each road-map. Our implementations of fundamental functions outperform the implementations on the same platform by 1.7× to 16.7×. Taking CKKS FHE schemes as a case study, our implementation of homomorphic multiplication achieves 3.2× speedup over the state-of-the-art GPU-based implementation, even considering the difference of platforms. The detailed evaluation and comparison of this paper would offer a vital reference for the follow-up work to choose appropriate underlying arithmetic units and important primitive optimizations in GPU-based FHE implementations.
Guang Fan 0001, Fangyu Zheng, Lipeng Wan 0002, Yuan Zhao 0015, Jiankuo Dong, Yuewu Wang, Jingqiang Lin 0001
IPDPS1
2022 A Novel High-Performance Implementation of CRYSTALS-Kyber with AI Accelerator
Lipeng Wan 0002, Fangyu Zheng, Guang Fan 0001, Rong Wei, Yuewu Wang, Jingqiang Lin 0001, Jiankuo Dong
ESORICS (3)3
2022 TEGRAS: An Efficient Tegra Embedded GPU-Based RSA Acceleration Server
abstract
Industrial Internet of Things (IIoT) has strict requirements on the performance and security of devices. Public-key cryptography, as a kind of computing resource-consuming algorithm, is widely used in the digital signature, key exchange, and so on. The embedded graphics processing units (GPUs) are now rapidly achieving extraordinary computing power, such as NVIDIA Tegra K1/X1/X2/Xavier, which are also treated as edge computing devices. They are widely used in IIoT environments, such as intelligent manufacturing, smart cities, and vehicle-mounted systems. The performance advantages endow embedded GPUs with the possibility of accelerating cryptography that also requires high-density computing. This article implements an efficient Tegra-based embedded GPU RSA acceleration server-oriented IIoT, named TEGRAS. Various optimization methods are employed to promote efficiency, including multithreaded Montgomery multiplication and Chinese Remainder Theorem implementation on the resource-constricted embedded GPUs. With about 40–50 W of power consumption, TEGRAS can deliver 34 kops/s of RSA2048 signature generation and 1007 kops/s of RSA signature verification, which outperforms implementations in the desktop GPUs and embedded CPUs in the perspective of performance-to-power ratio. To evaluate TEGRAS in real-world scenarios, we additionally build a network stack to deliver digital signature services, which can provide more than 34 and 978 kops of signature generation and signature verification, respectively. In a word, based on the embedded GPU, we provide a high-throughput, low-latency, and ready-to-use RSA accelerator-oriented IIoT.
Jiankuo Dong, Guang Fan 0001, Fangyu Zheng, Tianyu Mao, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Internet Things J.2
2022 EC-ECC: Accelerating Elliptic Curve Cryptography for Edge Computing on Embedded GPU TX2
abstract
Driven by artificial intelligence and computer vision industries, Graphics Processing Units (GPUs) are now rapidly achieving extraordinary computing power. In particular, the NVIDIA Tegra K1/X1/X2 embedded GPU platforms, which are also treated as edge computing devices, are now widely used in embedded environments such as mobile phones, game consoles, and vehicle-mounted systems to support high-dimension display, auto-pilot, and so on. Meanwhile, with the rise of the Internet of Things (IoT), the demand for cryptographic operations for secure communications and authentications between edge computing nodes and IoT devices is also expanding. In this contribution, instead of the conventional implementations based on FPGA, ASIC, and ARM CPUs, we provide an alternative solution for cryptographic implementation on embedded GPU devices. Targeting the new cipher suite added in TLS 1.3, we implement Edwards25519/448 and Curve25519/448 on an edge computing platform, embedded GPU NVIDIA Tegra X2, where various performance optimizations are customized for the target platform, including a novel parallel method for the register-limited embedded GPUs. With about 15 W of power consumption, it can provide 210k/31k ops/s of Curve25519/448 scalar multiplication, 834k/123k ops/s of fixed-point Edwards25519/448 scalar multiplication, and 150k/22k ops/s of unknown-point one, which are respectively the primitives and main workloads of key agreement, signature generation, and verification of the TLS 1.3 protocol. Our implementations achieve 8 to 26 times speedup of OpenSSL running in the very powerful ARM CPU of the same platform and outperform the state-of-the-art implementations in FPGA by a wide margin with better power efficiency.
Jiankuo Dong, Fangyu Zheng, Jingqiang Lin 0001, Zhe Liu 0001, Fu Xiao 0001, Guang Fan 0001
ACM Trans. Embed. Comput. Syst.6
2021 Heterogeneous-PAKE: Bridging the Gap between PAKE Protocols and Their Real-World Deployment
abstract
Two entities, who only share a password and communicate over an insecure channel, authenticate each other and agree on a large session key for protecting their subsequent communication. This is called the password-authenticated key exchange (PAKE) protocol. PAKE protocol has been considered a suitable substitute for the prevailing hash-based authentication which is vulnerable to various attacks. However, vendors are discouraged by both its prohibitively computational overheads as well as integrating costs, leading to its limited use since being proposed.
Rong Wei, Fangyu Zheng, Jiankuo Dong, Guang Fan 0001, Lipeng Wan 0002, Jingqiang Lin 0001, Yuewu Wang
ACSAC5
2021 TX-RSA: A High Performance RSA Implementation Scheme on NVIDIA Tegra X2
Jiankuo Dong, Guang Fan 0001, Fangyu Zheng, Jingqiang Lin 0001, Fu Xiao 0001
WASA (2)2
2021 SECCEG: A Secure and Efficient Cryptographic Co-processor Based on Embedded GPU System
Guang Fan 0001, Fangyu Zheng, Jiankuo Dong, Jingqiang Lin 0001, Rong Wei, Lipeng Wan 0002
WASA (2)1
2020 SEGIVE: A Practical Framework of Secure GPU Execution in Virtualization Environment
abstract
With the advancement of processor technology, general-purpose GPUs have become popular parallel computing accelerators in the cloud. However, designed for graphics rendering and high-performance computing, GPUs are born without sound security mechanisms. Consequently, the GPU-based service in the cloud is vulnerable to attacks from the potentially compromised guest OS as large amounts of sensitive code and data are offloaded directly to the unprotected GPUs.In this paper, we propose SEGIVE, a practical framework of secure GPU execution in the virtualization environment, which protects offloaded device code and data from disclosure or tampering by malicious guest OSes through the full life cycle of security-critical GPU applications. First, SEGIVE secures all the traffic transferred to GPUs with Intel SGX technology, including the users' sensitive data and GPU binaries. Second, with various memory isolation mechanisms, SEGIVE enhances security in multi-user execution scenarios by sharing a GPU among multiple workloads, which avoids underutilization of device resources. Besides, SEGIVE requires no modifications to application source codes, the GPU architecture, or I/O interconnection to fulfill security principles, and thus almost all prevailing GPU-based applications can easily benefit from SEGIVE with little porting effort. We have implemented SEGIVE with KVM-QEMU on off-the-shelf NVIDIA GPUs and CPUs. Evaluation results show that with security-enhances, the performance of SEGIVE prototype is still competitive to the native execution on compute-intensive applications, especially for the public-key cryptography algorithm.
Fangyu Zheng, Jingqiang Lin 0001, Guang Fan 0001, Jiankuo Dong
IPCCC4