VLDB 2026 Research / reviewers in the wild / expert
Fei Lyu 0002
dblp:86/7981-2
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-2282-1574ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 6 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph-Structure-Aware Hyperdimensional Computing for Hardware Trojan Detection
Zilong Su, Fei Lyu 0002, Yongjun Xia, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2026 | Optimized NTT Architecture Based on the Plantard Algorithm for ML-KEM and ML-DSAabstractModular multiplication is a vital operation in the Number Theoretic Transform (NTT), significantly enhancing polynomial multiplication in Post-Quantum Cryptography (PQC). The design efficiency of modular multiplication directly influences the computational performance of polynomial computation units. This work marks the first hardware-oriented improvement of the Plantard algorithm, optimizing the NTT architecture. We modify the Plantard algorithm and propose three innovative enhanced versions tailored for lattice-based cryptography (LBC). By employing pre-processed twiddle factors for result correction and eliminating an additional constant multiplication, we greatly simplify the computation steps. Based on these improvements, we further design a lightweight BRAM-free iterative NTT and a high-speed Multi-path Delay Commutator (MDC) pipelined NTT, both targeting the ML-KEM and ML-DSA parameter sets. Implementation results on the Xilinx Artix-7 platform demonstrate that our Plantard_preω design reduces slice usage by 22% to 53% and delay by 41.1% to 44.3% compared to existing algorithms like Barrett and K2RED. Furthermore, our iterative NTT design achieves the minimal area-time product (ATP) among state-of-the-art implementations, with reductions of 61.3% and 40.8% in ENS, and frequency increases of 72.7% and 145.4% for ML-KEM and ML-DSA, respectively. The pipelined NTT design also reduces delay by 23.7% and ATP by 4.9%, showcasing the compactness and superior performance of our approach. Bei Wang 0013, Ziying Ni, Mengxue Li, Fei Lyu 0002, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 5 |
| 2025 | An Efficient Methodology for Binary Logarithmic Computations of Floating-Point Numbers With Normalized Output Within One ulp of AccuracyabstractMany studies have focused on the hardware implementation of binary logarithmic computation with fixed-point output. Although their outputs are accurate within 1 ulp (unit in the last place) in fixed-point format, they are far from meeting the accuracy requirement of 1 ulp in floating-point format when the output is close to 0. However, normalized floating-point output that is accurate to within 1-3 ulp is needed in many math libraries (for example, OpenCL, NVIDIA CUDA, and AMD AOCL). To the best of our knowledge, this is the first study to propose a hardware implementation of binary logarithmic computation for floating-point numbers with a normalized output that is accurate to within 1 ulp. Instead of calculating$\textrm{log}_{2}(1+fi)$(where$\boldsymbol{fi}$is the fractional part of the floating-point number) directly, the proposed methodology uses two novel objective functions for the polynomial approximation method. The novel objective functions make the significant bits of the outputs move forward to eliminate the necessity for high precision near zero. Compared with the designs of fixed-point binary logarithmic converters, the proposed hardware implementation achieves greater accuracy to meet the requirement of 1 ulp of floating-point format with a 21% extra area consumption. Fei Lyu 0002, Yuanyong Luo, Weiqiang Liu 0001 |
IEEE Trans. Computers | 1 |
| 2025 | A Universal Methodology of Complex Number Computation for Low-Complexity and High-Speed ImplementationabstractIn complex-valued neural network (CVNN) applications, complex number calculations require high performance rather than high precision. However, most previous studies focused on high-precision approaches, which have low speed and high hardware costs. This paper proposes a universal methodology of complex number computation for low-complexity and high-speed implementation. The proposed methodology is based on the piecewise linear (PWL) method and can be used for different types of complex number computations. Considering that multiplication operations consume considerable resources, multiplication, fused square-add (FSA) and fused multiply-add (FMA) operations are the focus of optimization. The partial products of the square operation are reduced by folding and merging techniques because of their symmetry in the FSA operation. The partial products of the multiplication and FMA operations are reduced via Booth encoding. In addition, the partial products are further reduced by the proposed step-by-step truncation method. The proposed segmenter, which simulates the hardware implementation, automatically divides the nonlinear functions in the complex number computations into the smallest number of segments according to the required precision. The results show that the proposed approach improves performance and reduces hardware costs compared with the state-of-the-art methods for complex number calculations involving square roots, reciprocals and logarithms. Yu Wang 0161, Youlong Wu, Fei Lyu 0002, Yuanyong Luo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | An Optimized Architecture for Computing the Square Root of Complex NumbersabstractIn this paper, we propose an optimized architecture for computing the square root of complex numbers based on the piecewise linear (PWL) method. In the proposed design, the square-add and multiply-add operations are the focus of optimization. The symmetry of the partial products of the square operation is used to reduce the number of partial products. In addition, the least significant bits (LSBs) of the square-add and multiply-add are truncated to reduce the bits of partial products. According to the optimization of the hardware circuit, the simulation of the circuit in the segmentor is modified by introducing the simulation of the truncated fused square-add operation and truncated fused multiply-add operation. Experimental results show that the proposed optimization architecture has superiority in area, delay and power when compared with state-of-the-art designs. Yu Wang 0161, Fei Lyu 0002, Yuanyong Luo |
ISCAS | 6 |
| 2022 | Reconfigurable Multifunction Computing Unit Using an Universal Piecewise Linear MethodabstractComputing units for nonlinear complex functions are indispensable in deep neural network training processors. However, the existing computing units for nonlinear complex functions have low utilization efficiency and poor agreement and precision. In this article, we propose a multifunction computing unit for training deep neural networks by reusing computing resources based on a piecewise linear (PWL) method to improve computing density. Based on the state-of-the-art segmentor of PWL method, multiple nonlinear functions are divided into the fewest segments with the same bit width of computation. In hardware implementation, the reconfigurable technique is implemented on multiple functions while reusing computing resources including the multiplier and adder. The application-specific integrated circuit (ASIC) implementation results reveal that the architecture with reuse reduces the area by 44.50% and the power by 43.71% at the same frequency, when compared with the architecture without reuse. Fei Lyu 0002, Wenxiu Wang, Yuanyong Luo, Yu Wang 0161 |
ISCAS | 1 |
| 2022 | ML-PLAC: Multiplierless Piecewise Linear Approximation for Nonlinear Function EvaluationabstractIn this article, we propose a multiplierless piecewise linear (PWL) approximation computation (ML-PLAC) method for nonlinear unary functions. ML-PLAC seeks the minimum number of segments with the predefined fractional bit width and number of adders to satisfy the restriction on the maximum absolute error (MAE). Compared with the previous universal PWL approximation method, multiplication operations in the segmentor are replaced by a simulation of shift-and-add operations by reducing the fractional bit width of the slope of linear functions. Various numbers of segments are obtained by different predefined numbers of adders to balance the two numbers. In addition, adders are used to replace the multiplier in hardware architecture. The synthesized results prove that ML-PLAC has increased performance without any compromises. Compared with state-of-the-art methods, ML-PLAC saves 58.82% area, 38.16% delay, 60.31% power and 6.10%MAEwhen computing logarithmic functions; 51.51% area, 46.49% delay, and 46.95% power while maintaining the comparableMAEwhen computing antilogarithmic functions; 55.21% area, 25% delay, 61.21% power, and 37.30%MAEwhen computing hyperbolic tangent functions; 82.47% area, 60% delay, 77.51% power and 12.74%MAEwhen computing sigmoid functions; and 46.43% area, 31.43% delay, and 61.16% power while maintaining the sameMAEwhen computing softsign functions. Fei Lyu 0002, Zhelong Mao, Yanxu Wang, Yu Wang 0161, Yuanyong Luo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | High-Throughput Low-Latency Pipelined Divider for Single-Precision Floating-Point NumbersabstractIn this brief, we propose a fully pipelined divider for single-precision floating-point numbers based on a universal piecewise linear (PWL) approximation method and a modified Goldschmidt algorithm. The state-of-the-art universal PWL method uses a suitable number of segments and fractional bit widths to meet the requirement of the predefined maximum absolute error. Small multipliers are employed in the modified Goldschmidt algorithm. In the hardware implementation, the multipliers are optimized with the radix-4 and radix-8 booth encoding methods to reduce the number of partial products. In addition, the sum of the partial products and other data are calculated by a compressor and an adder to shorten the critical path. Synthesized results show that the maximum achievable frequency of our design is better than those of the existing methods. In addition, our design shows overwhelming superiority in terms of latency and throughput compared with existing methods. Fei Lyu 0002, Yanxu Wang, Yuanyong Luo, Yu Wang 0161 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Ultralow-Latency VLSI Architecture Based on a Linear Approximation Method for Computing Nth Roots of Floating-Point NumbersabstractState-of-the-art approaches that perform root computations based on the COordinate Rotation Digital Computer (CORDIC) algorithm suffer from high latency in performing multiple iterations. Therefore, root computations based on the CORDIC algorithm cannot meet the strict latency requirements of some applications. In this paper, we propose a methodology for performing Nth root computations on floating-point numbers based on the piecewise linear (PWL) approximation method. The proposed method divides an Nth root computation into several subtasks approximated by the PWL algorithm. It determines the widest segments of the subtasks and the smallest fractional width needed to satisfy the predefined maximum relative error Max_Errr. Our design is coded in Verilog HDL and synthesized under TSMC 40 nm CMOS technology. The synthesized results show that our design can reach the highest frequency of 2.703 GHz with an area consumption of 2608.84 μ m2and a power consumption of 2.4476 mW. Compared with one stateof-the-art architecture, our design saves 91.60%, 89.84%, and 63.33% of the area, power, and latency @1.89GHz frequency, respectively, while reducing Max_Errrby 57.30%. In addition, it saves 94.52%, 92.68%, and 73.17% of the area, power, and delay @1.89GHz frequency, respectively, and reduces Max_Errrby 1.65% when compared with the other state-of-the-art design. Fei Lyu 0002, Xiaoqi Xu, Yu Wang 0161, Yuanyong Luo, Hongbing Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | PWL-Based Architecture for the Logarithmic Computation of Floating-Point NumbersabstractIn this brief, we propose a logarithmic converter for floating-point numbers based on the piecewise linear (PWL) approximation method. The proposed method is applicable to any customized floating-point format with a mantissa length of 16–23 bits and a maximum absolute error (MAE) larger than 10−6. The logarithmic function is automatically segmented into several maximal subsections by a software-based segmentation scheme with the restriction of a predefined MAE and a fractional word length for the computing units. Then, we make a tradeoff between the piecewise number and the fractional word length. Based on the results of the segmentor, our design is coded in the Verilog hardware description language. The synthesized results show that our design consumes less area, time, and power without compromising accuracy compared to existing techniques based on the COordinate Rotation Digital Computer (CORDIC) and PWL methods. Fei Lyu 0002, Zhelong Mao, Yu Wang 0161, Yuanyong Luo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |