Xianglong Deng

dblp:231/1891 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
13since 2021 · last 2026
0009-0002-2058-5109ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption Accelerator
abstract
Fully homomorphic encryption (FHE) enables computation on encrypted data without compromising privacy, positioning it as a promising solution for secure cloud computing. However, its substantial computational overhead impedes practical deployment, prompting the development of dedicated hardware accelerators. In practice, when deploying cryptographic algorithm optimizations on FHE accelerators, hardware constraints typically such as limited memory capacity, often lead to a disparity between theoretical algorithmic advantage and achievable hardware efficiency.
Liang Kong 0005, Xianglong Deng, Guang Fan 0001, Shengyu Fan, Yilan Zhu, Geng Yang 0001, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005
ASPLOS (2)2
2026 An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design Automation
abstract
Fully Homomorphic Encryption (FHE) has become a promising approach to protecting data privacy in emerging application scenarios. Unfortunately, FHE suffers from significant processing speed degradation compared to plaintext computation, with one of the primary bottlenecks being the time-consuming Number Theoretic Transform (NTT). Therefore, accelerating NTT to accommodate various FHE parameters is crucial to advancing FHE towards practical use. With highly reconfigurable and performant logical fabrics, Field Programmable Gate Arrays (FPGAs) have exhibited great potential in NTT acceleration. By decomposing large-point NTT with strong data dependency into independent and simple small-point NTTs, the emerging Ten-step NTT (TNTT) algorithms intuitively enable higher parallelism and thereby have the potential to explore better performance compared to traditional algorithms. However, our quantitative analysis reveals that TNTT exhibits significant performance degradation as parallelism increases due to additional varying-size transpositions and Hadamard products. This paper proposes AutoNest, an efficient and scalable hardware architecture, along with an accelerator auto-generation framework for TNTT. The proposed hardware architecture maximizes performance by 1) adopting a 2D block decomposition dataflow to address critical path delays in transpose logic, thereby improving clock frequency. 2) integrating algorithm-level costfree twiddle factor fusion to reduce the number of modular multiplications in Hadamard products, thereby allowing higher parallelism on chip. Moreover, we also deliver an accelerator generation framework conducting automated design space exploration to elaborate a performant TNTT architecture under the target FPGAs' resource budget for user-defined FHE parameters. Experimental results on the AMD-Xilinx U280 FPGA demonstrate that NTT accelerators generated by AutoNest achieve an average speedup of$2.31 \times$compared to prior designs.
Yilan Zhu, Geng Yang 0001, Xingyu Tian, Dilshan Kumarathunga, Liang Kong 0005, Xianglong Deng, Shengyu Fan, Guang Fan 0001, Guiming Shi, Bo Zhang 0098, Yisong Chang, Shoumeng Yan, Zhenman Fang, Mingzhe Zhang 0005
HPCA6
2026 HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Geng Yang 0001, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng, Meng Li 0004, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005
ISCA10
2026 MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access Optimization
Xianglong Deng, Guang Fan 0001, Shengyu Fan, Mingzhe Zhang 0005
ISCA2
2026 A Bitwidth-Flexible Modular Multiplier with Shift-Free Accumulation for Efficient NTT Acceleration in FHE
Shengyu Fan, Xianglong Deng, Rui Hou 0001, Mingzhe Zhang 0005
ISCAS4
2025 The Future of Fully Homomorphic Encryption System: From a Storage I/O Perspective
Erci Xu, Shengyu Fan, Xianglong Deng, Guiming Shi, Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Shoumeng Yan, Mingzhe Zhang 0005
APPT5
2025 WPC: Weight Plaintext Compression for CNN Inference based on RNS-CKKS
abstract
Convolutional neural network (CNN) inference based on RNS-CKKS enables secure processing on encrypted data but introduces significant weight size overhead. Weight plaintext, weight in RNS-CKKS format, can reach tens to hundreds of gigabytes. Existing compression methods either add high computational cost or yield low compression rates. In this work, we propose WPC, Weight Plaintext Compression, to compress weight plaintext for RNS-CKKS-based CNN inference. We observe that the transformation from the weight in CNN models to the weight plaintext in RNS-CKKS format involves an operation akin to the Discrete Fourier Transform, which shifts data between the time and frequency domains while retaining redundant information from periodic and discrete data. Based on this observation, we first introduce the Periodic Transmit Theorem, which states that periodic patterns can be preserved during the transformation process, thereby enabling compression. We then propose Channel Innermost Packing Scheme and Rotation Padding to rearrange the weight data into periodic patterns for compression. Results show that WPC achieves 1.25 to 2.18 times speedup on an A100 GPU and 46.08 to 139.11 times compression rate.
Guiming Shi, Shengyu Fan, Xianglong Deng, Liang Kong 0005, Jingwei Cai, Shuwen Deng, Mingzhe Zhang 0005, Kaisheng Ma
CCS4
2025 WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA Cores
abstract
The application of Fully Homomorphic Encryption (FHE) is rapidly gaining traction as a means to maintain data confidentiality while performing computations on encrypted data. Given the accessibility and computational power, GPUs hold promise for significantly accelerating FHE operations. However, existing GPU-based acceleration solutions face several formidable challenges, notably the extensive occurrence of pipeline stalls induced by memory access and suboptimal harnessing of GPU hardware. This paper presents WarpDrive, a comprehensive framework for GPU-based FHE acceleration. Through sophisticated computation decomposition and fine-grained memory access design, WarpDrive significantly reduces the number of instructions by $\mathbf{7 3 \%}$ and pipeline stalls by $\mathbf{8 6 \%}$ compared to the state-of-the-art solution. Additionally, WarpDrive features a framework that supports the concurrent utilization of CUDA Cores and Tensor Cores within the NTT operation, for the first time, achieving performance that surpasses that of any single type of processing unit. Furthermore, we fully exploit the intra-ciphertext parallelism to elevate both computation and memory utilization, achieving up to $2.12 \times$ improvements without the need for ciphertext batching. Experimental results demonstrate that our optimizations highly enhance the performance of homomorphic operations. On an NVIDIA A100 GPU, WarpDrive achieves a throughput of 1218 KOPS for NTT and 305 KOPS for homomorphic multiplication, outperforming the state-of-the-art GPU solution (TensorFHE) by factors of $13.4 \times$ and $3.5 \times$, respectively. For the specific FHE workload, even under a much smaller batch size, our approach achieves $2.8 \times$ the performance of TensorFHE.
Guang Fan 0001, Mingzhe Zhang 0005, Fangyu Zheng, Shengyu Fan, Xianglong Deng, Wenxu Tang, Liang Kong 0005, Shoumeng Yan
HPCA6
2025 FAST: An FHE Accelerator for Scalable-parallelism with Tunable-bit
abstract
Fully Homomorphic Encryption (FHE) enables direct computation on encrypted data, providing substantial security advantages in cloud-based modern society.However, FHE suffers from significant computational overhead compared to plaintext computation, hindering its adoption in real-world applications.While many accelerators have been designed to address performance bottlenecks, most do not fully leverage cryptographic optimization technologies, leaving room for further performance enhancements.In this work, we propose FAST, an FHE accelerator incorporating recent cryptographic optimizations, including hoisting technology and the gadget decomposition key-switching method (named KLSS method).We analyze ciphertext level consumption throughout application execution and observe that workload requirements vary significantly with different ciphertext levels for both hybrid and KLSS key-switching methods.Additionally, we note the differing computational precision requirements for these key-switching methods.Based on these observations, we designed a versatile framework that supports multiple key-switching methods during a single application execution and integrates hoisting technology.
Shengyu Fan, Xianglong Deng, Liang Kong 0005, Guiming Shi, Guang Fan 0001, Dan Meng 0002, Rui Hou 0001, Mingzhe Zhang 0005
ISCA2
2025 Neo: Towards Efficient Fully Homomorphic Encryption Acceleration using Tensor Core
abstract
Fully Homomorphic Encryption (FHE) is an emerging cryptographic technique for privacy-preserving computation, which enables computations on the encrypted data.Nonetheless, the massive computational demands of FHE prevent its further application to real-world workloads.To tackle this problem, several studies focus on the ASIC-based acceleration for FHE.However, the rapid evolution of FHE algorithms poses challenges to the generality of ASIC accelerator design.By contrast, a number of works rely on GPGPUs for FHE accelerations, due to the high parallelism and flexibility provided by GPGPUs.In this work, we propose a GPGPU-based acceleration solution that supports the Cheon-Kim-Kim-Song (CKKS) scheme by further exploiting Tensor Core(TCU) capabilities.In our study, we * Both author contributed equally to this research.
Xianglong Deng, Shengyu Fan, Dan Meng 0002, Rui Hou 0001, Mingzhe Zhang 0005
ISCA2
2025 HAWK: Fully Homomorphic Encryption Accelerator with Fixed-Word Key Decomposition Switching
Liang Kong 0005, Shengyu Fan, Xianglong Deng, Guang Fan 0001, Guiming Shi, Yilan Zhu, Geng Yang 0001, Shoumeng Yan, Mingzhe Zhang 0005
MICRO3
2025 LP-HENN: fully homomorphic encryption accelerator with high energy efficiency
abstract
Abstract Fully homomorphic encryption (FHE) enables direct computation on encrypted data without decryption, ensuring data privacy in cloud computing scenarios and preventing the leakage of sensitive information. However, the computational overhead of HE typically exceeds that of plaintext computation by 4 to 5 orders of magnitude, while energy consumption is 5 to 6 orders of magnitude higher. These substantial performance and energy overheads significantly hinder the widespread adoption of FHE. This paper proposed LP-HENN, a novel low-power and energy-efficient FHE accelerator architecture that leverages a RISC-V vector coprocessor and ReRAM crossbar arrays. LP-HENN targets power-constrained application scenarios such as edge devices, aiming to provide highly energy-efficient acceleration support for FHE applications. LP-HENN leverages the collaborative work of the vector processor and ReRAM crossbars, employing optimization strategies to achieve full pipelining and minimize memory access. Furthermore, this paper proposed a parameter selection model for early-stage architecture design, which achieves an optimal balance between performance and energy consumption through the collaborative optimization of multiple parameters. Experimental results show that, for an FHE-based convolutional neural network (HE-CNN) inference application, LP-HENN achieves a 31.82Ã- and 11920.56Ã- improvement in performance and energy efficiency, respectively, compared to CPU. Compared to FxHENN, the state-of-the-art FPGA-based FHE accelerator with high energy efficiency for edge devices, LP-HENN achieves a 2.36Ã- and 10.04Ã- improvement in performance and energy efficiency, respectively. The energy efficiency of LP-HENN is comparable to that of F1, the state-of-the-art ASIC FHE accelerator, while featuring a low power design suitable for edge computing.
Zhuoyu Tian, Shengyu Fan, Xianglong Deng, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005
Cybersecur.4
2024 Trinity: A General Purpose FHE Accelerator
abstract
Fully Homomorphic Encryption (FHE) is crucial for privacy-preserving computing, which allows direct computation on encrypted data. While various FHE schemes have been proposed, none of them efficiently support both arithmetic FHE and logic FHE simultaneously. To address this issue, researchers explore the combination of different FHE schemes within a single application and propose algorithms for the conversion between them. Unfortunately, all prior ASIC-based FHE accelerators are designed to support a single FHE scheme, and none of them supports the acceleration for FHE scheme conversion. This necessitates FHE acceleration systems to integrate multiple accelerators for different schemes, leading to increased system complexity and hindering performance enhancement. In this paper, we present the first multi-modal FHE accelerator based on a unified architecture, which efficiently supports CKKS, TFHE, and their conversion scheme within a single accelerator. To achieve this goal, we first analyze the theoretical foundations of the aforementioned schemes and highlight their composition from a finite number of arithmetic kernels. Then, we investigate the challenges for efficiently supporting these kernels within a unified architecture, which include 1) concurrent support for NTT and FFT, 2) maintaining high hardware utilization across various polynomial lengths, and 3) ensuring consistent performance across diverse arithmetic kernels. To tackle these challenges, we propose a novel FHE accelerator named Trinity, which in-corporates algorithm optimizations, hardware component reuse, and dynamic workload scheduling to enhance the acceleration of CKKS, TFHE, and their conversion scheme. By adaptive select the proper allocation of components for NTT and MAC, Trinity maintains high utilization across NTTs with various polynomial lengths and imbalanced arithmetic workloads. The experiment results show that, for the pure CKKS and TFHE workloads, the performance of our Trinity outperforms the state-of-the- art accelerator for CKKS (SHARP) and TFHE (Morphling) by 1.49 x and 4.23 x, respectively. Moreover, Trinity achieves 919.3 x performance improvement for the FHE-conversion scheme over the CPU-based implementation. Notably, despite the performance improvement, the hardware overhead of Trinity is only 85 % of the summed circuit areas of SHARP and Morphling.
Xianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian, Jiangrui Yu, Dingyuan Cao 0002, Dan Meng 0002, Rui Hou 0001, Meng Li 0004, Qian Lou, Mingzhe Zhang 0005
MICRO1
2020 Superpixel-Level Weighted Label Propagation for Hyperspectral Image Classification
abstract
As a typical graph-based semisupervised learning technique, the label propagation (LP) approach has gained much attention in recent years. The key to LP algorithms is the propagation capability and efficiency of the similarity matrix, which describes the similarity between two data points. Concerning hyperspectral image which often contains hundreds of thousands of pixels, the corresponding similarity matrix is particularly huge and thus the LP procedure is intractable. Fortunately, superpixel, which can effectively characterize the spatial semantic information of surface objects, provides a reasonable way to solve this problem. In this article, we propose an elaborate superpixel-based weighted LP approach, abbreviated as SuWLP, for hyperspectral image classification. First, the hyperspectral image is oversegmented by the entropy rate segmentation (ERS) method, and the internal consistency of each superpixel can be achieved. Second, a new similarity measure is designed to estimate the similarity between two superpixels, and a superpixel-based similarity matrix can be thus established. Third, after the training samples have been expanded based on the superpixel distribution, a weighted LP technique is designed to propagate the sample label at the superpixel level without any parameter tuning. Finally, the label of each superpixel maps back to the contained pixels. We compared our proposed SuWLP method with several state-of-the-art ones, and experimental results on three real hyperspectral data sets certify the effectiveness and efficiency of the superpixel-level LP strategy.
Sen Jia 0001, Xianglong Deng, Meng Xu 0002, Jun Zhou 0001, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.2
2019 Collaborative Representation-Based Multiscale Superpixel Fusion for Hyperspectral Image Classification
abstract
In virtue of the spatial structural characteristic of surface materials, the performance of the hyperspectral image classification can be boosted by incorporating texture information. Normally, the spatial structure can be extracted by predefined operators, including the popular extended multiattribute profiles (EMAPs) and the Gabor filters. Recently, superpixel segmentation, which reflects the homogeneous regularity of objects, has drawn much attention in the field. In this paper, a collaborative representation-based multiscale superpixel fusion (CRMSF) approach has been proposed for the hyperspectral image classification. First, after obtaining the EMAPs from the raw hyperspectral image, a group of predesigned 3-D Gabor wavelet filters is convolved with the EMAP features, and the EMAP-Gabor features can, thus, be achieved. Second, the collaborative representation-based classification (CRC) is employed to fully and efficiently make use of the huge amount of extracted EMAP-Gabor features. Third, multiscale superpixel maps are generated from the EMAP features that are utilized to regularize the classification map obtained by CRC. A heuristic strategy has been specially devised to automatically decide the number of extracted superpixels in multiple scales, which can be perfectly compatible with hyperspectral images having various spatial sizes and spatial resolutions. This is the most important contribution of the developed CRMSF approach. Finally, the classification task is accomplished by fusing the multiple regularized classification maps. The CRMSF approach has been evaluated on four popular hyperspectral image data sets, and the experimental results show the advantages of CRMSF, particularly for a hyperspectral image with high spatial resolution.
Sen Jia 0001, Xianglong Deng, Jiasong Zhu, Meng Xu 0002, Jun Zhou 0001, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.2
2018 Extended Morphological Profile-based Gabor Wavelets for Hyperspectral Image Classification
abstract
Hyperspectral image acquired by a hyperspectral sensor contains hundred of narrow contiguous spectral bands, since the spatial distribution of surface materials generally exhibits high regularity and local continuity, spatial texture information should be introduced to improve the classification accuracy of hyperspectral image. The extended morphological profiles (EMP) have been created from the raw hyperspectral image, which has proven to be effective and robust of reflecting the spatial structural features of hyperspectral data. In the meanwhile, because the three-dimensional (3D) Gabor wavelets have been introduced to exploit the joint spectral-spatial features of hyperspectral image. In this paper, for the purpose of combining the advantages of the EMP operator and Gabor wavelet transform together, an extended morphological profile-based Gabor wavelets, which is named as EMP-Gabor, has been proposed for hyperspectral image classification. Definitely, to compute principal components of the hyperspectral image, the most significant principal components are used as base images for an extended morphological profile, and the EMP features can be thus obtained. Secondly, 3D Gabor wavelets with particular orientations are directly convolved with the EMP feature cube. Finally, support vector machine (SVM) classifier is utilized to carry out the classification task. Experimental results on two real hyperspectral data sets have demonstrated the effectiveness of the proposed EMP-Gabor framework for hyperspectral image classification over several state-of-the-art methods.
Sen Jia 0001, Huimin Xie, Xianglong Deng
ICPR3