Zhihan Xu

dblp:89/4436 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HERA: A Bandwidth-efficient Accelerator for Fully Homomorphic Encryption on HBM-enabled FPGA
abstract
Fully Homomorphic Encryption (FHE) enables privacy-preserving computation on encrypted data. However, it incurs massive computation and DRAM traffic overheads, making hardware acceleration essential. Existing FPGA-based solutions offer limited exploration of bandwidth utilization and memory optimizations, leaving room for further performance improvements.
Zhihan Xu, Rajgopal Kannan, Viktor Prasanna 0001
FPGA1
2026 Function-Aligned Nonuniform Quantization Graph Coloring for Distributed Functional Compression
Zhihan Xu, Juan Alberto Cabrera Guerrero, Frank H. P. Fitzek
ICC1
2026 TSUIE: Efficient two-stage underwater image enhancement framework
Lingfeng Chen, Tianheng Ma, Yuanxin Xu, Zhihan Xu, Shibo Lu, Yunong Wang
Neurocomputing4
2025 FAST: FPGA Acceleration of Fully Homomorphic Encryption with Efficient Bootstrapping
abstract
Bootstrapping is a critical operation in Fully Homomorphic Encryption (FHE) for privacy-preserving computation. Due to its significant computational overhead, accelerating bootstrapping is crucial for practical FHE applications involving deep evaluation circuits. In this paper, we introduce FAST, an FPGA-based accelerator for efficient FHE bootstrapping. We propose novel datapath optimizations for two key operations in bootstrapping: homomorphic linear transformation (HLT) and polynomial evaluation. Our memory-efficient datapath designed for HLT significantly reduces off-chip ciphertext access. We also speed up the polynomial evaluation process by reducing the number of required HE operations. We conduct an in-depth analysis of the Advanced Bootstrapping Algorithm (ABA) and highlight its computational advantages. FAST is the first accelerator to support ABA, demonstrating significant speedup for bootstrapping. In addition, we develop a novel versatile permutation circuit to handle diverse permutation patterns in FHE, achieving high throughput and efficient resource utilization. Compared with the state-of-the-art (SOTA) GPU and FPGA designs, FAST achieves 8.84× and 5.89× speedups for bootstrapping, respectively. As illustrative examples of deep FHE applications, we show that FAST delivers over 20× speedup for logistic regression training compared with the SOTA GPU implementation and outperforms the SOTA FPGA design by 1.43× for ResNet-20 inference.
Zhihan Xu, Tian Ye 0002, Rajgopal Kannan, Viktor Prasanna 0001
FPGA1
2025 FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
abstract
Homomorphic Encryption (HE) enables secure computation on encrypted data, addressing privacy concerns in cloud computing. However, the high computational cost of HE operations, particularly matrix multiplication (MM), remains a major barrier to its practical deployment. Accelerating Homomorphic Encrypted MM (HE MM) is crucial for applications such as privacy-preserving machine learning. In this paper, we present a bandwidth-efficient FPGA implementation of HE MM. We first develop a cost model to evaluate the on-chip memory requirement for a given set of HE parameters and input matrix sizes. Our analysis shows that optimizing on-chip memory usage is critical for scalable and efficient HE MM. To this end, we design a novel datapath for Homomorphic Linear Transformation (HLT), the major bottleneck in HE MM. Our datapath significantly reduces off-chip memory traffic and on-chip memory demand by enabling fine-grained data reuse. Leveraging the proposed datapath, we introduce FAME, the first FPGA-based accelerator specifically tailored for HE MM. FAME supports arbitrary matrix shapes and is configurable for a wide range of HE parameter sets. We implement FAME on Alveo U280 and evaluate its performance over diverse matrix sizes and shapes. Experimental results show that FAME achieves an average of$221 \times$speedup over state-of-the-art CPU-based implementations, demonstrating its scalability and practicality for large-scale consecutive HE MM and real-world workloads.
Zhihan Xu, Rajgopal Kannan, Viktor Prasanna 0001
FPL1
2025 BDMUIE: Underwater image enhancement based on Bayesian diffusion model
Lingfeng Chen, Zhihan Xu, Yuanxin Xu
Neurocomputing2
2024 Tracing the Evolution of Information Transparency for OpenAI's GPT Models through a Biographical Approach
abstract
Information transparency, the open disclosure of information about models, is crucial for proactively evaluating the potential societal harm of large language models (LLMs) and developing effective risk mitigation measures. Adapting the biographies of artifacts and practices (BOAP) method from science and technology studies, this study analyzes the evolution of information transparency within OpenAI’s Generative Pre-trained Transformers (GPT) model reports and usage policies from its inception in 2018 to GPT-4, one of today’s most capable LLMs. To assess the breadth and depth of transparency practices, we develop a 9-dimensional, 3-level analytical framework to evaluate the comprehensiveness and accessibility of information disclosed to various stakeholders. Findings suggest that while model limitations and downstream usages are increasingly clarified, model development processes have become more opaque. Transparency remains minimal in certain aspects, such as model explainability and real-world evidence of LLM impacts, and the discussions on safety measures such as technical interventions and regulation pipelines lack in-depth details. The findings emphasize the need for enhanced transparency to foster accountability and ensure responsible technological innovations.
Zhihan Xu, Eni Mustafaraj
AIES (1)1
2024 Bandwidth Efficient Homomorphic Encrypted Discrete Fourier Transform Acceleration on FPGA
abstract
Fully Homomorphic Encryption (FHE) plays an important role in privacy-preserving computation on the cloud. It allows computations on encrypted data without decryption. Bootstrapping is a fundamental operation in FHE, enabling an unlimited number of homomorphic encrypted computations, but at a significant time cost. A major bootstrapping component, the Homomorphic Encrypted Discrete Fourier Transform (HE DFT), is particularly time-consuming and requires the transfer of a large amount of data from external memory. In this paper, we propose a bandwidth-efficient FPGA implementation of HE DFT. We design a cost model to evaluate the on-chip memory requirement and the off-chip data transfer overhead for HE DFT. Our analysis shows that prior approaches can lead to significant off-chip data transfers, which process the entire ciphertext between subroutines. To address DRAM transfer overhead, we propose LimbFlow, an optimized dataflow approach for HE DFT that enhances fine-grained data reuse by rearranging the processing order of ciphertext and merging several subroutines. Leveraging the LimbFlow, we develop an FPGA-based accelerator tailored for HE DFT. We evaluate the accelerator on AMD U280 FPGA across various sets of security parameters. Our accelerator achieves up to 4.90 × and 1.98 × speedup compared with the State-Of-The-Art (SOTA) GPU and FPGA implementations.
Zhihan Xu, Yang Yang 0111, Rajgopal Kannan, Viktor Prasanna 0001
FCCM1
2022 N3H-Core: Neuron-designed Neural Network Accelerator via FPGA-based Heterogeneous Computing Cores
abstract
Accelerating the neural network inference by FPGA has emerged as a popular option, since the reconfigurability and high performance computing capability of FPGA intrinsically satisfies the computation demand of the fast-evolving neural algorithms. However, the popular neural accelerators on FPGA (e.g., Xilinx DPU) mainly utilize the DSP resources for constructing their processing units, while the rich LUT resources are not well exploited. Via the software-hardware co-design approach, in this work, we develop an FPGA-based heterogeneous computing system for neural network acceleration. From the hardware perspective, the proposed accelerator consists of DSP- and LUT-based GEneral Matrix-Multiplication (GEMM) computing cores, which forms the entire computing system in a heterogeneous fashion. The DSP- and LUT-based GEMM cores are computed w.r.t a unified Instruction Set Architecture (ISA) and unified buffers. Along the data flow of the neural network inference path, the computation of the convolution/fully-connected layer is split into two portions, handled by the DSP- and LUT-based GEMM cores asynchronously. From the software perspective, we mathematically and systematically model the latency and resource utilization of the proposed heterogeneous accelerator, regarding varying system design configurations. Through leveraging the reinforcement learning technique, we construct a framework to achieve end-to-end selection and optimization of the design specification of target heterogeneous accelerator, including workload split strategy, mixed-precision quantization scheme, and resource allocation of DSP- and LUT-core. In virtue of the proposed design framework and heterogeneous computing system, our design outperforms the state-of-the-art Mix&Match design with latency reduced by 1.12-1.32x with higher inference accuracy. The N3H-core is open-sourced at: https://github.com/elliothe/N3H_Core.
Zhihan Xu, Zhezhi He, Weifeng Zhang 0003, Xiaobing Tu, Xiaoyao Liang, Li Jiang 0002
FPGA2
2004 Design of a novel knowledge-based fault detection and isolation scheme
abstract
In this paper, a real-time fault detection and isolation (FDI) scheme for dynamical systems is developed, by integrating the signal processing technique with neural network design. Wavelet analysis is applied to capture the fault-induced transients of the measured signals in real-time, and the decomposed signals are pre-processed to extract details about a fault. A Regional Self-Organizing feature Map (R-SOM) neural network is synthesized to classify the fault types. The R-SOM neural network adopts two regions adjustment in the learning algorithm, thus it has high precision in clustering and matching, especially when the noise, disturbance and other uncertainties exist in the systems. As a result, the proposed FDI scheme is robust and accurate. The design is implemented on a stirred tank system and satisfactory online testing results are obtained.
Qing Zhao 0003, Zhihan Xu
IEEE Trans. Syst. Man Cybern. Part B2