EDBT 2026 Demo / reviewers in the wild / expert
Xiao Hu 0007
dblp:19/1374-7
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0001-7668-4689ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ALT: Area-Efficient and Low-Latency FPGA Design for Torus Fully Homomorphic EncryptionabstractThe homomorphic encryption over the torus (TFHE) is a promising fully homomorphic encryption (FHE) scheme that allows arbitrary homomorphic computations with the programmable bootstrapping (PBS) algorithm. However, PBS suffers from prohibitive computation complexity and latency, which hinders the practical applications of TFHE. To address these challenges, we propose ALT, a field-programmable gate array (FPGA) accelerator for PBS that exhibits high area efficiency and low latency. Our approach involves modifying the parameters of the PBS algorithm to strike a balance between the computation complexity and the decryption failure rate (DFR). In addition, we leverage the Chinese residue theorem (CRT) to exploit the inherent parallelism and construct the primes to eliminate the need of CRT process and facilitate fast modular arithmetic. The ALT design comprises several carefully designed computation units, including inverse CRT (ICRT), divide-and-round (DR) operation, and monomial number theoretic transform (MNTT). We employ algorithmic and architectural co-optimization techniques to optimize these units. Notably, ALT features a low-complexity MNTT module, enabling the utilization of the bootstrapping key unrolling (BKU) technique with reduced latency and minimal hardware resources. Furthermore, all submodules of ALT are parameterized and scalable, allowing the entire design to be configurable according to varying requirements across different application scenarios. Experimental results on FPGA demonstrate that ALT significantly outperforms a similar configurable work in terms of latency, throughput, and efficiency. In comparison with the fastest FPGA implementation, ALT can realize lower latency while reducing digital signal processor (DSP) reduction by over$50\%$, leading to enhanced area efficiency and energy efficiency. Xiao Hu 0007, Zhihao Li 0001, Zhongfeng Wang 0001, Xianhui Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Reconfigurable and High-Efficiency Polynomial Multiplication Accelerator for CRYSTALS-KyberabstractRecently, the National Institute of Standards and Technology (NIST) has identified the first four quantum-resistant algorithms for post-quantum cryptography (PQC) standardization. CRYSTALS-Kyber (Kyber) is the only public-key encryption and key-establishment algorithm among them. In this article, we propose a reconfigurable, high-speed, and area-efficient polynomial multiplication accelerator for Kyber to facilitate its practical applications. The cornerstone of polynomial multiplication is the butterfly unit (BU) structure, composed of modular addition, subtraction, and multiplication. For the modular multiplication, we adopt the Barrett reduction method and reduce the size of operands leveraging the form of modulus with a novel formula transformation, which significantly decreases the computational complexity and increases the maximum clock frequency. On the hardware side, we make four BU modules constitute a binomial arithmetic core (Bi-Core) as the basic reconfigurable unit. The memory access scheme tailored for parallel processing is explored with data-reusing and memory-grouping methods, and a compact control logic is devised. The complete polynomial multiplication architecture is coded with Verilog and implemented on a Xilinx Artix-7 xc7a100t-3 device. Experiment results demonstrate that our implementations with different configurations all outperform the state-of-the-art works in area efficiency by up to 39% improvement in terms of area-time product (ATP). Moreover, the proposed design with four Bi-Cores achieves the fastest speed among existing designs. Minghao Li 0001, Jing Tian 0004, Xiao Hu 0007, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | AC-PM: An Area-Efficient and Configurable Polynomial Multiplier for Lattice Based CryptographyabstractAs the computation bottleneck in lattice-based cryptography (LBC), the polynomial multiplication based on number theoretic transform (NTT) has been continuously studied for flexible hardware implementations with high area-efficiency. This paper presents an area-efficient and configurable NTT-based polynomial multiplier (AC-PM) incorporating algorithmic and architectural level optimization techniques. For the core operation of polynomial multiplication, two low-complexity and fast modular multiplication algorithms are introduced with loose constraints of LBC-friendly primes. Based on the proposed algorithms, a reconfigurable processing element (RPE) is dedicatedly designed to execute all the operations in an NTT-based polynomial multiplication: NTT, inverse NTT (INTT), and coefficient-wise multiplication (CWM). The proposed AC-PM can be configured with different numbers of RPEs and supports various polynomial degrees without recompilation. Additionally, the dataflow complexity is greatly simplified. More importantly, to the best of our knowledge, the twiddle factors are reused, for the first time, to support both NTT and INTT with multiple polynomial degrees, which leads to increased flexibility of AC-PM with small overhead on hardware resource. FPGA implementation results demonstrate that the proposed AC-PM significantly outperforms the prior arts in both flexibility and area efficiency. Xiao Hu 0007, Jing Tian 0004, Minghao Li 0001, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | A High-Speed FPGA-Based Hardware Implementation for Leighton-Micali SignatureabstractDue to the rapid progress made in quantum computers, modern cryptography faces great challenges. Many digital signature schemes that have resistance to quantum computing are studied and standardized by several influential international organizations. The Leighton-Micali signature (LMS) protocol, one of the hash-based signature schemes, is standardized by both the Internet Engineering Task Force (IETF) and the National Institute of Standards and Technology (NIST) due to its well-studied security and relatively small signature size. However, the heavy computation load and high latency of LMS limits its practical applications. In this paper, for the first time, we propose a full hardware implementation of LMS to accelerate all the three stages:$key~generation$,$signature~generation$, and$verification$. Considering the scalability requirement and the characteristic of the parameter sets of LMS, we extract the coarse-grained basic logic, a hash group, and build a reconfigurable architecture for all available parameters by carefully designing the parallelism degree while achieving low latency and high hardware utilization efficiency. Then, we devise a fusion architecture for$key~generation$and$signature~generation$based on the hash group module. Moreover, for the$signature~verification$stage, we propose a separate architecture by applying the hash group module along with an efficient depth-first Merkle tree module. We code our designs with Verilog language in parameterized style and implement them on a Xilinx XCVU7P FPGA platform. The experimental results show that significant improvements are obtained for different parameter sets by the proposed designs when compared to state-of-the-art works. Yifeng Song, Xiao Hu 0007, Jing Tian 0004, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Efficient Homomorphic Convolution Designs on FPGA for Secure InferenceabstractRecently, secure neural network (NN) inference, a combination of homomorphic encryption (HE) and NN, has attracted much attention. Nevertheless, a large number of computations, mainly brought by the HE scheme, form the bottleneck in real-time applications. In this article, we present a hardware accelerator on a field-programmable gate array (FPGA) for the homomorphic convolution layer (HomConvL), which is the most computation-intensive part of the HE-based secure inference. First, we propose a new HomConvL algorithm called packed rotations at inputs (PaRotI), which is suitable for hardware implementation for its inherent high parallelism and low complexity with acceptable noise growth and moderate resource consumption. Then, we present three highly parallel architectures for different parameter sets and application scenarios of state-of-the-art HomConvL algorithms. The new architectures are implemented on a Xilinx VCU110 FPGA board, and the experimental results demonstrate that our designs can achieve 15.31–$19.46\times $speedups compared with the software implementations. Xiao Hu 0007, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | DARM: A Low-Complexity and Fast Modular Multiplier for Lattice-Based CryptographyabstractThe lattice-based cryptography (LBC) has been widely used recently in many compute-intensive applications, such as the post-quantum cryptography (PQC) and privacy-preserving deep learning, where the main task for such applications is to improve the computational efficiency. The modular multiplication operations, mainly involved in the number theoretic transform (NTT), comprise a large proportion of the whole computations required by an LBC. This paper presents a novel "decompose-and-reduce" modular multiplication algorithm (DARM), considering primes with the form of q = 22N−δ and δN−2. The inherent structure of the modulus is exploited and the intermediates’ data widths are reduced. Moreover, a low-complexity and fast multiplier is elaborately devised based on DARM. To further validate the performance of our multiplier, an n-point NTT design with DARM is implemented with various configurations. FPGA implementation results demonstrate that compared with the prior arts, the proposed multiplier has 1.12-1.89× speedups with the least DSP utilization. For the case of ⌈log2q⌉ = 60 and n = 4096, the NTT implementation with DARM achieves up to 41.2% and 61.2% reductions in LUTs and DSPs, respectively. Xiao Hu 0007, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001 |
ASAP | 1 |
| 2021 | High-Speed and Scalable FPGA Implementation of the Key Generation for the Leighton-Micali Signature ProtocolabstractDue to the rapid progress made in quantum computers, modern cryptography faces great challenges. Many new digital signature schemes that have resistance to quantum computing are being presented for Post-Quantum Cryptography (PQC) standardization. The Leighton-Micali signature (LMS), a kind of hash-based signature scheme, is selected as a promising candidate for the PQC signature protocols by the Internet Engineering Task Force (IETF) because of its small private and public key sizes. However, the low-efficiency in key generation forms the bottleneck in practical applications. In this paper, we propose a high-speed architecture for the key generation to accelerate the LMS for the first time. The architecture is delicately devised to be scalable, supporting all the parameter sets for the LMS. The degree of parallelism is carefully designed to achieve low latency and high hardware utilization efficiency. Moreover, the control flow is well managed to accommodate different parameter sets with constant power for the consideration of anti-power analysis attacks. We code our design with Verilog language and implement it on the Xilinx Zynq UltraScale+ FPGA. The experimental results show that, compared with the optimal software implementation running on an Intel(R) Core(TM) i7-6850K 3.60GHz CPU with threading enabled, the new design achieves 55x to 2091x speedup in different parameter configurations. Yifeng Song, Xiao Hu 0007, Jing Tian 0004, Zhongfeng Wang 0001 |
ISCAS | 2 |