Minghao Li 0001

dblp:91/1271-1 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-5490-2024ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2026 High-Throughput and Configurable Modular Multiplier Using Extended Montgomery Reduction
Yueqin Dai, Minghao Li 0001, Zhongfeng Wang 0001
ISCAS2
2026 Rethinking Area Optimization for High-Throughput Hardware Accelerator of Number Theoretic Transform: A Case Study on ML-KEM
abstract
The computation-intensive modular multiplications in number theoretic transform (NTT) present significant bottlenecks for lattice-based post-quantum cryptography (PQC). In this article, we propose a systematic area optimization strategy through joint algorithm-hardware codesign. First, we propose two high-throughput architectures to accommodate diverse throughput requirements: a fully parallel NTT and a folded 64-point variant derived from the former, which both feature area-efficient constant modular multipliers. A novel permutation-based algorithm is developed to reuse the NTT architecture for inverse NTT without structural modifications. Second, an area-driven, automated decomposition-based design strategy for constant modular multipliers is developed to reduce hardware overhead. The strategy involves a three-stage process: decomposing, analyzing, and selecting, which are based on finite-field properties and comprehensive circuit modeling. Finally, the proposed architectures are evaluated on Artix-7 AC701 FPGA with the parameters of the module-lattice-Based key-encapsulation mechanism (ML-KEM) from the latest PQC standard. The results show that the optimized modular multipliers deliver an average resource reduction of approximately 46% over the prior art. Furthermore, operating at 189 MHz, the core throughputs reach 47.2 and$94.5\times 10^{6}$operations per second (OPS) for the proposed folded and fully parallel structures, respectively, while both architectures achieve an I/O-limited throughput of$23.6{\,}\times{\,}10^{6}$OPS. Compared to the state-of-the-art highly parallel designs, our work provides a 31.1% area efficiency improvement.
Minghao Li 0001, Suwen Song, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2025 HRCIM-NTT: An Efficient Compute-in-Memory NTT Accelerator With Hybrid-Redundant Numbers
abstract
Recently, four NIST-approved Post-Quantum Cryptography (PQC) algorithms are selected to be standardized. Three of them are lattice-based cryptographic schemes and feature the number-theoretic transform (NTT) as the computing bottleneck compelling fast and low-power hardware implementations. In this work, a high-speed and power-efficient NTT accelerator is presented leveraging the compute-in-memory (CIM) technique with bottom-up optimizations. Firstly, a carry-free modular multiplication (CFMM) algorithm is proposed, which utilizes on-the-fly reduction and hybrid-redundant representation to optimize the butterfly unit operation, the cornerstone of NTT. Based on the optimized algorithm, an efficient butterfly unit in memory (BUIM) is developed by co-designing with SRAM circuit, which saves the memory access energy, decreases operation cycles, and obtains ultra-short critical path. Additionally, the data pattern of CIM array is also improved to avoid redundant memory read/write operations, which further reduces memory access overhead. Finally, a combination of pipelined operation flow and constant interstage data mapping strategy is employed to bestow the proposed hybrid-redundant CIM NTT (HRCIM-NTT) architecture with minimized computing cycles and reduced routing overhead. The implementation under 45nm CMOS technology demonstrates that HRCIM-NTT achieves the highest throughput and lowest latency among the existing CIM-based NTT accelerators.
Xu Zhang 0040, Yaodong Wei, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 Low-latency Hardware Architecture for VDF Evaluation in Class Groups
abstract
The verifiable delay function (VDF), as a kind of cryptographic primitives, has recently been adopted quite often in decentralized systems. Highly correlated to the security of VDFs, the fastest implementation for VDF evaluation is generally desired to be publicly known. In this paper, for the first time, we propose a low-latency hardware implementation for the complete VDF evaluation in the class group by jointly exploiting optimizations. On one side, we reduce the required computational cycles by decreasing the hardware-unfriendly divisions and increase the parallelism of computations by reducing the data dependency. On the other side, we provide low-latency large-number divisors, multipliers, and adders, respectively, while those operators are generally very hard to be accelerated. Besides, we carefully schedule the sub-modules and devise the low-latency architecture for the complete VDF evaluation. Finally, the proposed design is coded and synthesized under the TSMC 28-nm CMOS technology. The experimental results show that our design can achieve a speedup of 3.5x compared to the optimal C++ implementation for the VDF evaluation over an advanced CPU. Moreover, compared to the state-of-the-art hardware implementation for the squaring, a key step of VDF, we achieve about 2x speedup.
Danyang Zhu, Jing Tian 0004, Minghao Li 0001, Zhongfeng Wang 0001
IEEE Trans. Computers3
2023 Reconfigurable and High-Efficiency Polynomial Multiplication Accelerator for CRYSTALS-Kyber
abstract
Recently, the National Institute of Standards and Technology (NIST) has identified the first four quantum-resistant algorithms for post-quantum cryptography (PQC) standardization. CRYSTALS-Kyber (Kyber) is the only public-key encryption and key-establishment algorithm among them. In this article, we propose a reconfigurable, high-speed, and area-efficient polynomial multiplication accelerator for Kyber to facilitate its practical applications. The cornerstone of polynomial multiplication is the butterfly unit (BU) structure, composed of modular addition, subtraction, and multiplication. For the modular multiplication, we adopt the Barrett reduction method and reduce the size of operands leveraging the form of modulus with a novel formula transformation, which significantly decreases the computational complexity and increases the maximum clock frequency. On the hardware side, we make four BU modules constitute a binomial arithmetic core (Bi-Core) as the basic reconfigurable unit. The memory access scheme tailored for parallel processing is explored with data-reusing and memory-grouping methods, and a compact control logic is devised. The complete polynomial multiplication architecture is coded with Verilog and implemented on a Xilinx Artix-7 xc7a100t-3 device. Experiment results demonstrate that our implementations with different configurations all outperform the state-of-the-art works in area efficiency by up to 39% improvement in terms of area-time product (ATP). Moreover, the proposed design with four Bi-Cores achieves the fastest speed among existing designs.
Minghao Li 0001, Jing Tian 0004, Xiao Hu 0007, Zhongfeng Wang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 AC-PM: An Area-Efficient and Configurable Polynomial Multiplier for Lattice Based Cryptography
abstract
As the computation bottleneck in lattice-based cryptography (LBC), the polynomial multiplication based on number theoretic transform (NTT) has been continuously studied for flexible hardware implementations with high area-efficiency. This paper presents an area-efficient and configurable NTT-based polynomial multiplier (AC-PM) incorporating algorithmic and architectural level optimization techniques. For the core operation of polynomial multiplication, two low-complexity and fast modular multiplication algorithms are introduced with loose constraints of LBC-friendly primes. Based on the proposed algorithms, a reconfigurable processing element (RPE) is dedicatedly designed to execute all the operations in an NTT-based polynomial multiplication: NTT, inverse NTT (INTT), and coefficient-wise multiplication (CWM). The proposed AC-PM can be configured with different numbers of RPEs and supports various polynomial degrees without recompilation. Additionally, the dataflow complexity is greatly simplified. More importantly, to the best of our knowledge, the twiddle factors are reused, for the first time, to support both NTT and INTT with multiple polynomial degrees, which leads to increased flexibility of AC-PM with small overhead on hardware resource. FPGA implementation results demonstrate that the proposed AC-PM significantly outperforms the prior arts in both flexibility and area efficiency.
Xiao Hu 0007, Jing Tian 0004, Minghao Li 0001, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Efficient Homomorphic Convolution Designs on FPGA for Secure Inference
abstract
Recently, secure neural network (NN) inference, a combination of homomorphic encryption (HE) and NN, has attracted much attention. Nevertheless, a large number of computations, mainly brought by the HE scheme, form the bottleneck in real-time applications. In this article, we present a hardware accelerator on a field-programmable gate array (FPGA) for the homomorphic convolution layer (HomConvL), which is the most computation-intensive part of the HE-based secure inference. First, we propose a new HomConvL algorithm called packed rotations at inputs (PaRotI), which is suitable for hardware implementation for its inherent high parallelism and low complexity with acceptable noise growth and moderate resource consumption. Then, we present three highly parallel architectures for different parameter sets and application scenarios of state-of-the-art HomConvL algorithms. The new architectures are implemented on a Xilinx VCU110 FPGA board, and the experimental results demonstrate that our designs can achieve 15.31–$19.46\times $speedups compared with the software implementations.
Xiao Hu 0007, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2021 DARM: A Low-Complexity and Fast Modular Multiplier for Lattice-Based Cryptography
abstract
The lattice-based cryptography (LBC) has been widely used recently in many compute-intensive applications, such as the post-quantum cryptography (PQC) and privacy-preserving deep learning, where the main task for such applications is to improve the computational efficiency. The modular multiplication operations, mainly involved in the number theoretic transform (NTT), comprise a large proportion of the whole computations required by an LBC. This paper presents a novel "decompose-and-reduce" modular multiplication algorithm (DARM), considering primes with the form of q = 22N−δ and δN−2. The inherent structure of the modulus is exploited and the intermediates’ data widths are reduced. Moreover, a low-complexity and fast multiplier is elaborately devised based on DARM. To further validate the performance of our multiplier, an n-point NTT design with DARM is implemented with various configurations. FPGA implementation results demonstrate that compared with the prior arts, the proposed multiplier has 1.12-1.89× speedups with the least DSP utilization. For the case of ⌈log2q⌉ = 60 and n = 4096, the NTT implementation with DARM achieves up to 41.2% and 61.2% reductions in LUTs and DSPs, respectively.
Xiao Hu 0007, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001
ASAP2
2020 A Three-Level Scoring System for Fast Similarity Evaluation Based on Smith-Waterman Algorithm
abstract
The Smith-Waterman (S-W) algorithm is widely adopted by the state-of-the-art DNA sequence aligners in next-generation sequencing (NGS). Prevailing read aligners, such as BWA-MEM and Bowtie 2, use the S-W algorithm to implement the seed-and-extend paradigm. In this work, we further extend the functionality of the S-W algorithm to evaluate the similarity between a pair of sequences without going through traceback process, and design a three-level hardware scoring system to compute final result efficiently. The system is made reconfigurable to align pairs of sequences of various length with a restriction of maximum number of errors. Experimental results show that the system can achieve a throughput of 685Mb/s at 69 iterations in the case of 126bp and the accuracy rate of the outputs is over 98% campared with software results. To the best of our knowledge, this is the first hardware implementation for a similarity evaluation system based on the S-W algorithm.
Jiajun Wu 0025, Minghao Li 0001, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS3