EDBT 2026 Demo / reviewers in the wild / expert
Zhihao Li 0001
dblp:40/2903-1
· DBLP profile ↗
20ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-6149-0627ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 5 since 2021Security and privacy · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sub-Millisecond Gate BootstrappingabstractGate bootstrapping is a core primitive that enables arbitrary circuit evaluation in fully homomorphic encryption (FHE), where blind rotation remains the dominant performance bottleneck. In this work, we present a sub-millisecond NTRU-based gate bootstrapping scheme that achieves state-of-the-art performance through coordinated algorithmic, software, and hardware-level optimizations. Chunling Chen, Zhihao Li 0001, Qingyun Niu, Xianhui Lu, Ruida Wang, Lutan Zhao, Rui Hou 0001 |
AsiaCCS | 2 |
| 2025 | Gibbon: Faster Secure Two-party Training of Gradient Boosting Decision TreeabstractGradient Boosting Decision Tree (GBDT) and its variants are widely used in industry. They have achieved remarkable success in numerous machine learning competitions and practical applications. Secure Multi-Party Computation (MPC) allows multiple data owners to compute a function jointly while keeping their input private. In this work, we present Gibbon, a secure two-party GBDT training framework on a vertically split dataset, where two data owners each hold different features of the same data samples. Compared with the state-of-the-art Squirrel (USENIX'Sec 2023), for most parameter settings, Gibbon achieves 2×-4× reduction in running time and 2×-3× reduction in communication. Lichun Li, Zecheng Wu, Yuan Zhao 0015, Zhihao Li 0001 |
CCS | 4 |
| 2025 | Full domain functional bootstrapping using the prime cyclotomic ring
Ruida Wang, Xianhui Lu, Yundi Wen, Zhihao Li 0001, Benqiang Wei, Kunpeng Wang 0001, Lixia Luo |
Theor. Comput. Sci. | 4 |
| 2025 | Chameleon: An Efficient FHE Scheme Switching Acceleration on GPUsabstractFully homomorphic encryption (FHE) enables direct computation on encrypted data, making it a crucial technology for privacy protection. However, FHE suffers from significant performance bottlenecks. In this context, GPU acceleration offers a promising solution to bridge the performance gap. Existing efforts primarily focus on single-class FHE schemes, which fail to meet the diverse requirements of data types and functions, prompting the development of hybrid multi-class FHE schemes. However, studies have yet to thoroughly investigate specific GPU optimizations for hybrid FHE schemes. In this paper, we present an efficient GPU-based FHE scheme switching acceleration named Chameleon. First, we propose a scalable NTT acceleration design that adapts to larger CKKS polynomials and smaller TFHE polynomials. Specifically, Chameleon tackles synchronization issues by fusing stages to reduce synchronization, employing polynomial coefficient shuffling to minimize synchronization scale, and utilizing an SM-aware combination strategy to identify the optimal switching point. Second, Chameleon is the first to comprehensively analyze and optimize critical switching operations. It introduces CMux-level parallelization to accelerate LUT evaluation and a homomorphic rotation-free matrixvector multiplication to improve repacking efficiency. Finally, Chameleon outperforms the state-of-the-art GPU implementations by 1.23× in CKKS HMUL and 1.15× in bootstrapping. It also achieves up to 4.87× and 1.51× speedups for TFHE bootstrapping compared to CPU and GPU versions, respectively, and delivers a 67.3× average speedup for scheme switching over CPU-based implementation. Haoqi He, Lutan Zhao, Peinan Li, Zhihao Li 0001, Dan Meng 0002, Rui Hou 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | TFHE Bootstrapping: Faster, Smaller and Time-Space Trade-Offs
Ruida Wang, Benqiang Wei, Zhihao Li 0001, Xianhui Lu, Kunpeng Wang 0001 |
ACISP (1) | 3 |
| 2024 | Circuit Bootstrapping: Faster and Smaller
Ruida Wang, Yundi Wen, Zhihao Li 0001, Xianhui Lu, Benqiang Wei, Kunpeng Wang 0001 |
EUROCRYPT (2) | 3 |
| 2024 | Efficient Blind Rotation in FHEW Using Refined Decomposition and NTT
Ying Liu 0078, Zhihao Li 0001, Ruida Wang, Xianhui Lu, Kunpeng Wang 0001 |
ISC (1) | 2 |
| 2024 | ALT: Area-Efficient and Low-Latency FPGA Design for Torus Fully Homomorphic EncryptionabstractThe homomorphic encryption over the torus (TFHE) is a promising fully homomorphic encryption (FHE) scheme that allows arbitrary homomorphic computations with the programmable bootstrapping (PBS) algorithm. However, PBS suffers from prohibitive computation complexity and latency, which hinders the practical applications of TFHE. To address these challenges, we propose ALT, a field-programmable gate array (FPGA) accelerator for PBS that exhibits high area efficiency and low latency. Our approach involves modifying the parameters of the PBS algorithm to strike a balance between the computation complexity and the decryption failure rate (DFR). In addition, we leverage the Chinese residue theorem (CRT) to exploit the inherent parallelism and construct the primes to eliminate the need of CRT process and facilitate fast modular arithmetic. The ALT design comprises several carefully designed computation units, including inverse CRT (ICRT), divide-and-round (DR) operation, and monomial number theoretic transform (MNTT). We employ algorithmic and architectural co-optimization techniques to optimize these units. Notably, ALT features a low-complexity MNTT module, enabling the utilization of the bootstrapping key unrolling (BKU) technique with reduced latency and minimal hardware resources. Furthermore, all submodules of ALT are parameterized and scalable, allowing the entire design to be configurable according to varying requirements across different application scenarios. Experimental results on FPGA demonstrate that ALT significantly outperforms a similar configurable work in terms of latency, throughput, and efficiency. In comparison with the fastest FPGA implementation, ALT can realize lower latency while reducing digital signal processor (DSP) reduction by over$50\%$, leading to enhanced area efficiency and energy efficiency. Xiao Hu 0007, Zhihao Li 0001, Zhongfeng Wang 0001, Xianhui Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Full Domain Functional Bootstrapping with Least Significant Bit Encoding
Zhihao Li 0001, Benqiang Wei, Ruida Wang, Xianhui Lu, Kunpeng Wang 0001 |
Inscrypt (1) | 1 |
| 2023 | OpenFFT: An Adaptive Tuning Framework for 3D FFT on ARM Multicore CPUsabstractThe sophisticated hierarchy and shared characteristics of cache in multicore CPU architectures bring challenges to the performance improvement of fundamental algorithms, especially in implementing and optimizing 3D FFT. 3D FFT is a memory-bounded algorithm that contains many highly discretized memory accesses. With the working set scaling, the data locality becomes poor, which is prone to cause serious memory access overhead, especially for high-dimensional data transposition. This paper proposes a 3D FFT optimization framework named OpenFFT. This framework optimizes the memory access of 3D FFT by the following methods, including 1) A novel tiling algorithm, Z-OpenFFT, based on the column-order algorithm for high-dimensional vectorization to improve data locality and eliminate transposition; 2) An efficient search algorithm Section-cache-aware algorithm to optimize the memory access of butterfly network of 1D FFT; 3) A multi-thread allocation model by analyzing the characteristics of cache hierarchy and task size to allocate threads adaptively. Experiments demonstrate that OpenFFT could obtain a more competitive performance than the best configuration of FFTW and ARMPL on ARM CPUs. Tun Chen, Haipeng Jia, Yunquan Zhang, Kun Li 0016, Zhihao Li 0001, Jianyu Yao, Chendi Li |
ICS | 5 |
| 2023 | Fregata: Faster Homomorphic Evaluation of AES via TFHE
Benqiang Wei, Ruida Wang, Zhihao Li 0001, Qinju Liu, Xianhui Lu |
ISC | 3 |
| 2023 | Generating Fast FFT Kernels on CPUs via FFT-Specific IntrinsicsabstractThis paper proposes an algorithm-specific instruction (ASI)-based fast Fourier transform (FFT) code generation framework, named FFTASI, to generate unified architecture independent butterfly kernels that can be transformed into architecture-dependent kernels by establishing the mapping between ASIs and architecture-specific instructions for various hardware platforms. FFTASI strikes a good balance between performance and productivity on CPUs. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Yuyan Sun, Yiwei Zhang 0009, Tun Chen |
PPoPP | 1 |
| 2023 | HE-Booster: An Efficient Polynomial Arithmetic Acceleration on GPUs for Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) enables secure offloading of computations to untrusted cloud servers as it allows computing on encrypted data. However, existing well-known FHE schemes suffer from heavy performance overheads. Thus numerous accelerations based on FPGAs, ASICs, and GPUs have been proposed. Compared to FPGAs and ASICs, GPUs have obvious advantages in productivity and development costs. And also, GPUs have already been widely deployed in commercial cloud or supercomputing centers. Therefore, we present HE-Booster, an efficient GPU-based FHE acceleration design. For single-GPU acceleration, a thorough systematic design is exploited to map five common phases in typical FHE schemes to the GPU parallel architecture. In particular, inspired by the regular architecture of NTT/INTT, a novel inter-thread local synchronization is proposed to exploit thread-level parallelism. For multi-GPU acceleration, we propose a scalable parallelization design that exploitsdata-level parallelismthrough fine-grained data partition under different representations. Finally, experiments on 1 NVIDIA GPU demonstrate that our work outperforms 251.7×, 78.5× and 164.9× than three mainstream CPU-based libraries HElib, SEAL, and PALISADE, and up to 170.5× speedup is obtained compared to the GPU-accelerated library cuHE. What's more, performing 8 homomorphic multiplications on 8 GPUs can deliver up to a 7.66× performance boost compared to a single-GPU implementation. Peinan Li, Rui Hou 0001, Zhihao Li 0001, Jiangfeng Cao, XiaoFeng Wang 0001, Dan Meng 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Automatic Generation of High-Performance FFT Kernels on Arm and X86 CPUsabstractThis article presents AutoFFT, a template-based code generation framework that can automatically generate high-performance FFT kernels for all natural-number radices. AutoFFT is based on the Cooley-Tukey FFT algorithm, which exploits the symmetric and periodic properties of the DFT matrix, as the outer parallelization framework. Because butterflies are the core operations of the Cooley-Tukey algorithm, we explore additional symmetric and periodic properties of the DFT matrix and formulate multiple optimized calculation templates to further reduce the number of floating-point operations for butterflies of arbitrary natural numbers. To fully exploit hardware resources, we encapsulate a series of optimizations in an assembly template optimizer. Given any DFT problem, AutoFFT automatically generates C FFT kernels using these calculation templates and converts them into efficient assembly kernels using the template optimizer. Through a series of experiments on Arm, Intel, and AMD processors, we show that AutoFFT-generated kernels can outperform those in Fastest Fourier Transform in the West (FFTW), the Arm Performance Libraries (ARMPL), and the Intel Math Kernel Library (MKL). Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Tun Chen, Richard W. Vuduc |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | AutoFFT: a template-based FFT codes auto-generation framework for ARM and X86 CPUsabstractThe discrete Fourier transform (DFT) is widely used in scientific and engineering computation. This paper proposes a template-based code generation framework named AutoFFT that can automatically generate high-performance fast Fourier transform (FFT) codes. AutoFFT employs the Cooley-Tukey FFT algorithm, which exploits the symmetric and periodic properties of the DFT matrix as the outer parallelization framework. To further reduce the number of floating-point operations of butterflies, we explore more symmetric and periodic properties of the DFT matrix and formulate two optimized calculation templates for prime and power-of-two radices. To fully exploit hardware resources, we encapsulate a series of optimizations in an assembly template optimizer. Given any DFT problem, AutoFFT automatically generates C FFT kernels using these two templates and transfers them to efficient assembly codes using the template optimizer. Experiments show that AutoFFT outperforms FFTW, ARMPL, and Intel MKL on average across all FFT types on ARMv8 and Intel x86-64 processors. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Tun Chen, Luning Cao |
SC | 1 |
| 2019 | Efficient parallel optimizations of a high-performance SIFT on GPUs
Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Shice Liu, Shigang Li 0002 |
J. Parallel Distributed Comput. | 1 |
| 2018 | Implementation and Optimization of Multi-dimensional Real FFT on ARMv8 Platform
Haipeng Jia, Zhihao Li 0001, Yunquan Zhang |
ICA3PP (2) | 3 |
| 2017 | HartSift: A High-Accuracy and Real-Time SIFT Based on GPUabstractScale Invariant Feature Transform (SIFT) is one of the most popular and robust feature extraction algorithms for its invariance to scale, rotation and illumination. It has been widely adopted in many fields, such as video tracking, image stitching, simultaneous localization and mapping (SLAM), structure from motion (SFM) and so on. However, high computational complexity constrains its further application in real-time systems. These systems have to make a tradeoff between accuracy and performance to achieve real-time feature extraction. They adopt other faster algorithms but with less accuracy, like SURF and PCA-SIFT. In order to address this problem, this paper proposes a GPU-accelerated SIFT using CUDA, named HartSift, which realizes high-accuracy and real-time feature extraction by making full use of computing resources of CPU and GPU within a single machine. Experiments show that, on the NIVDIA GTX TITAN Black GPU, HartSift can process an image within 3.14?10.57ms (94.61?318.47fps) according to the size of images. In addition, HartSift is 59.34?75.96 times and 4.01?6.49 times faster than OpenCV-SIFT (a CPU version) and SiftGPU (a GPU version), respectively. In the mean time, HartSift's performance and CudaSIFT's (the fastest GPU version so far) are almost the same, while HartSift's accuracy is much higher than CudaSIFT's. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang |
ICPADS | 1 |
| 2015 | Causal discovery on high dimensional data
Ruichu Cai, Wen Wen 0009, Zhihao Li 0001 |
Appl. Intell. | 5 |
| 2014 | A Causal Model for Disease Pathway Discovery
Ruichu Cai, Chang Yuan, Wen Wen 0009, Zhihao Li 0001 |
ICONIP (1) | 7 |