Yuanyong Luo

dblp:228/8905 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-4450-066XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 2 first-author · 8 since 2021
YearPublicationVenuePosition
2025 An Efficient Methodology for Binary Logarithmic Computations of Floating-Point Numbers With Normalized Output Within One ulp of Accuracy
abstract
Many studies have focused on the hardware implementation of binary logarithmic computation with fixed-point output. Although their outputs are accurate within 1 ulp (unit in the last place) in fixed-point format, they are far from meeting the accuracy requirement of 1 ulp in floating-point format when the output is close to 0. However, normalized floating-point output that is accurate to within 1-3 ulp is needed in many math libraries (for example, OpenCL, NVIDIA CUDA, and AMD AOCL). To the best of our knowledge, this is the first study to propose a hardware implementation of binary logarithmic computation for floating-point numbers with a normalized output that is accurate to within 1 ulp. Instead of calculating$\textrm{log}_{2}(1+fi)$(where$\boldsymbol{fi}$is the fractional part of the floating-point number) directly, the proposed methodology uses two novel objective functions for the polynomial approximation method. The novel objective functions make the significant bits of the outputs move forward to eliminate the necessity for high precision near zero. Compared with the designs of fixed-point binary logarithmic converters, the proposed hardware implementation achieves greater accuracy to meet the requirement of 1 ulp of floating-point format with a 21% extra area consumption.
Fei Lyu 0002, Yuanyong Luo, Weiqiang Liu 0001
IEEE Trans. Computers2
2025 A Universal Methodology of Complex Number Computation for Low-Complexity and High-Speed Implementation
abstract
In complex-valued neural network (CVNN) applications, complex number calculations require high performance rather than high precision. However, most previous studies focused on high-precision approaches, which have low speed and high hardware costs. This paper proposes a universal methodology of complex number computation for low-complexity and high-speed implementation. The proposed methodology is based on the piecewise linear (PWL) method and can be used for different types of complex number computations. Considering that multiplication operations consume considerable resources, multiplication, fused square-add (FSA) and fused multiply-add (FMA) operations are the focus of optimization. The partial products of the square operation are reduced by folding and merging techniques because of their symmetry in the FSA operation. The partial products of the multiplication and FMA operations are reduced via Booth encoding. In addition, the partial products are further reduced by the proposed step-by-step truncation method. The proposed segmenter, which simulates the hardware implementation, automatically divides the nonlinear functions in the complex number computations into the smallest number of segments according to the required precision. The results show that the proposed approach improves performance and reduces hardware costs compared with the state-of-the-art methods for complex number calculations involving square roots, reciprocals and logarithms.
Yu Wang 0161, Youlong Wu, Fei Lyu 0002, Yuanyong Luo
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 An Optimized Architecture for Computing the Square Root of Complex Numbers
abstract
In this paper, we propose an optimized architecture for computing the square root of complex numbers based on the piecewise linear (PWL) method. In the proposed design, the square-add and multiply-add operations are the focus of optimization. The symmetry of the partial products of the square operation is used to reduce the number of partial products. In addition, the least significant bits (LSBs) of the square-add and multiply-add are truncated to reduce the bits of partial products. According to the optimization of the hardware circuit, the simulation of the circuit in the segmentor is modified by introducing the simulation of the truncated fused square-add operation and truncated fused multiply-add operation. Experimental results show that the proposed optimization architecture has superiority in area, delay and power when compared with state-of-the-art designs.
Yu Wang 0161, Fei Lyu 0002, Yuanyong Luo
ISCAS7
2022 Reconfigurable Multifunction Computing Unit Using an Universal Piecewise Linear Method
abstract
Computing units for nonlinear complex functions are indispensable in deep neural network training processors. However, the existing computing units for nonlinear complex functions have low utilization efficiency and poor agreement and precision. In this article, we propose a multifunction computing unit for training deep neural networks by reusing computing resources based on a piecewise linear (PWL) method to improve computing density. Based on the state-of-the-art segmentor of PWL method, multiple nonlinear functions are divided into the fewest segments with the same bit width of computation. In hardware implementation, the reconfigurable technique is implemented on multiple functions while reusing computing resources including the multiplier and adder. The application-specific integrated circuit (ASIC) implementation results reveal that the architecture with reuse reduces the area by 44.50% and the power by 43.71% at the same frequency, when compared with the architecture without reuse.
Fei Lyu 0002, Wenxiu Wang, Yuanyong Luo, Yu Wang 0161
ISCAS5
2022 ML-PLAC: Multiplierless Piecewise Linear Approximation for Nonlinear Function Evaluation
abstract
In this article, we propose a multiplierless piecewise linear (PWL) approximation computation (ML-PLAC) method for nonlinear unary functions. ML-PLAC seeks the minimum number of segments with the predefined fractional bit width and number of adders to satisfy the restriction on the maximum absolute error (MAE). Compared with the previous universal PWL approximation method, multiplication operations in the segmentor are replaced by a simulation of shift-and-add operations by reducing the fractional bit width of the slope of linear functions. Various numbers of segments are obtained by different predefined numbers of adders to balance the two numbers. In addition, adders are used to replace the multiplier in hardware architecture. The synthesized results prove that ML-PLAC has increased performance without any compromises. Compared with state-of-the-art methods, ML-PLAC saves 58.82% area, 38.16% delay, 60.31% power and 6.10%MAEwhen computing logarithmic functions; 51.51% area, 46.49% delay, and 46.95% power while maintaining the comparableMAEwhen computing antilogarithmic functions; 55.21% area, 25% delay, 61.21% power, and 37.30%MAEwhen computing hyperbolic tangent functions; 82.47% area, 60% delay, 77.51% power and 12.74%MAEwhen computing sigmoid functions; and 46.43% area, 31.43% delay, and 61.16% power while maintaining the sameMAEwhen computing softsign functions.
Fei Lyu 0002, Zhelong Mao, Yanxu Wang, Yu Wang 0161, Yuanyong Luo
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 High-Throughput Low-Latency Pipelined Divider for Single-Precision Floating-Point Numbers
abstract
In this brief, we propose a fully pipelined divider for single-precision floating-point numbers based on a universal piecewise linear (PWL) approximation method and a modified Goldschmidt algorithm. The state-of-the-art universal PWL method uses a suitable number of segments and fractional bit widths to meet the requirement of the predefined maximum absolute error. Small multipliers are employed in the modified Goldschmidt algorithm. In the hardware implementation, the multipliers are optimized with the radix-4 and radix-8 booth encoding methods to reduce the number of partial products. In addition, the sum of the partial products and other data are calculated by a compressor and an adder to shorten the critical path. Synthesized results show that the maximum achievable frequency of our design is better than those of the existing methods. In addition, our design shows overwhelming superiority in terms of latency and throughput compared with existing methods.
Fei Lyu 0002, Yanxu Wang, Yuanyong Luo, Yu Wang 0161
IEEE Trans. Very Large Scale Integr. Syst.5
2021 Ultralow-Latency VLSI Architecture Based on a Linear Approximation Method for Computing Nth Roots of Floating-Point Numbers
abstract
State-of-the-art approaches that perform root computations based on the COordinate Rotation Digital Computer (CORDIC) algorithm suffer from high latency in performing multiple iterations. Therefore, root computations based on the CORDIC algorithm cannot meet the strict latency requirements of some applications. In this paper, we propose a methodology for performing Nth root computations on floating-point numbers based on the piecewise linear (PWL) approximation method. The proposed method divides an Nth root computation into several subtasks approximated by the PWL algorithm. It determines the widest segments of the subtasks and the smallest fractional width needed to satisfy the predefined maximum relative error Max_Errr. Our design is coded in Verilog HDL and synthesized under TSMC 40 nm CMOS technology. The synthesized results show that our design can reach the highest frequency of 2.703 GHz with an area consumption of 2608.84 μ m2and a power consumption of 2.4476 mW. Compared with one stateof-the-art architecture, our design saves 91.60%, 89.84%, and 63.33% of the area, power, and latency @1.89GHz frequency, respectively, while reducing Max_Errrby 57.30%. In addition, it saves 94.52%, 92.68%, and 73.17% of the area, power, and delay @1.89GHz frequency, respectively, and reduces Max_Errrby 1.65% when compared with the other state-of-the-art design.
Fei Lyu 0002, Xiaoqi Xu, Yu Wang 0161, Yuanyong Luo, Hongbing Pan
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 PWL-Based Architecture for the Logarithmic Computation of Floating-Point Numbers
abstract
In this brief, we propose a logarithmic converter for floating-point numbers based on the piecewise linear (PWL) approximation method. The proposed method is applicable to any customized floating-point format with a mantissa length of 16–23 bits and a maximum absolute error (MAE) larger than 10−6. The logarithmic function is automatically segmented into several maximal subsections by a software-based segmentation scheme with the restriction of a predefined MAE and a fractional word length for the computing units. Then, we make a tradeoff between the piecewise number and the fractional word length. Based on the results of the segmentor, our design is coded in the Verilog hardware description language. The synthesized results show that our design consumes less area, time, and power without compromising accuracy compared to existing techniques based on the COordinate Rotation Digital Computer (CORDIC) and PWL methods.
Fei Lyu 0002, Zhelong Mao, Yu Wang 0161, Yuanyong Luo
IEEE Trans. Very Large Scale Integr. Syst.5
2020 A CORDIC-Based Architecture with Adjustable Precision and Flexible Scalability to Implement Sigmoid and Tanh Functions
abstract
In the artificial neural networks, tanh (hyperbolic tangent) and sigmoid functions are widely used as activation functions. Past methods to compute them may have shortcomings such as low precision or inflexible architecture that is difficult to expand, so we propose a CORDIC-based architecture to implement sigmoid and tanh functions, which has adjustable precision and flexible scalability. It just needs shift-add-or-subtract operations to compute high-accuracy results and is easy to expand the input range through scaling the negative iterations of CORDIC without changing the original architecture. We adopt the control variable method to explore the accuracy distribution through software simulation. A specific case (ARCH. (1, 15, 18), RMSE: 10−6) is designed and synthesized under the TSMC 40nm CMOS technology, the report shows that it has the area of 36512.78μm2and power of 12.35mW at the frequency of 1GHz. The maximum work frequency can reach 1.5GHz, which is better than the state-of-the-art methods.
Hui Chen 0015, Yuanyong Luo, Zhonghai Lu, Li Li 0003, Zongguang Yu
ISCAS3
2020 An Optimized Compression Strategy for Compressor-Based Approximate Multiplier
abstract
Approximate multipliers have recently attracted great attention due to their substantially lower energy consumption and area overhead. But previous approximate multiplier designs are mainly focused on the design of approximate compressors, little attention is paid to the compression strategy of partial product matrix. This paper proposes an optimized universal compression scheme for the compressor-based approximate multiplier. When we apply the new compression scheme to the state-of-the-art compressors, the accuracy of the approximate multiplier is largely increased and fewer exact adders are needed. To prove the efficiency of the new compression strategy, an 8-bit and a 12-bit approximate multipliers are designed using Verilog and synthesized under the TSMC 40-nm CMOS technology. Compared to the state-of-the-art, the experimental results indicate that the mean error distance of 8-bit multiplier decreases by 19.6%, with area and power reduced by 5.38% and 2.38% respectively; 12-bit multiplier has a reduction of 18.1% for mean error distance, with area and power reduced by 6.29% and 3.24% respectively. Moreover, application to image processing is presented, which shows that the proposed approximate multiplier has a better performance.
Manzhen Wang, Yuanyong Luo, Mengyu An, Yuou Qiu, Muhan Zheng, Zhongfeng Wang 0001, Hongbing Pan
ISCAS2
2020 PLAC: Piecewise Linear Approximation Computation for All Nonlinear Unary Functions
abstract
This article presents a piecewise linear approximation computation (PLAC) method for all nonlinear unary functions, which is an enhanced universal and error-flattened piecewise linear (PWL) approximation approach. Compared with the previous methods, PLAC features two main parts, an optimized segmenter to seek the minimum number of segments under the predefined software maximum absolute error (MAE), raising the segmentation performance to the highest theoretical level for logarithm, and a novel quantizer to completely simulate the hardware behavior and determine the required bit width and MAEc(MAE in circuits) for hardware implementation. In addition, the hardware architecture is also improved by simplifying the indexing logic, leading to nonredundant hardware overhead. The ASIC implementation results reveal that the proposed PLAC can improve all metrics without any compromise. Compared with the state-of-the-art methods, when computing logarithmic function, PLAC reduces 2.80% area, 3.77% power consumption, and 1.83% MAEcwith the same delay; when approximating hyperbolic tangent function, PLAC reduces 6.25% area, 4.31% power consumption, and 18.86% MAEcwith the same delay; when evaluating sigmoid function, PLAC reduces 16.50% area, 4.78% power consumption with the same delay, and MAEc; and when calculating softsign function, PLAC reduces 17.28% area, 11.34% power consumption, 12.50% delay, and 33.28% MAEc.
Hongxi Dong, Manzhen Wang, Yuanyong Luo, Muhan Zheng, Mengyu An, Yajun Ha, Hongbing Pan
IEEE Trans. Very Large Scale Integr. Syst.3
2020 GH CORDIC-Based Architecture for Computing $N$ th Root of Single-Precision Floating-Point Number
abstract
This article presents hardware implementation for computing arbitrary roots of a single-precision floating-point number. The proposed architecture is based on Generalized Hyperbolic COordinate Rotation Digital Computer (GH CORDIC) algorithm. Benefiting from the wide range of floating-point numbers, our design is able to compute the Nth root (N ≥ 2) of a single-precision floating-point number. After implementation, a series of tests have been carried out, including accuracy, power consumption, performance comparison, and so on. Simulation results indicate that our proposed method is capable of calculating the Nth root of a positive single-precision floating-point number with a relative error of 10-7approximately and promises an error-flatten performance. Synthesized results from a design compiler under TSMC-40-nm CMOS technology show that our design can achieve the highest frequency of 2.38 GHz with the area consumption of 140894.44 μm2and power consumption of 86.9573 mW.
Yuanyong Luo, Zhongfeng Wang 0001, Qinghong Shen, Hongbing Pan
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Generalized Hyperbolic CORDIC and Its Logarithmic and Exponential Computation With Arbitrary Fixed Base
abstract
This paper proposes a generalized hyperbolic COordinate Rotation Digital Computer (GH CORDIC) to directly compute logarithms and exponentials with an arbitrary fixed base. In a hardware implementation, it is more efficient than the state of the art which requires both a hyperbolic CORDIC and a constant multiplier. More specifically, we develop the theory of GH CORDIC by adding a new parameter called base to the conventional hyperbolic CORDIC. This new parameter can be used to specify the base with respect to the computation of logarithms and exponentials. As a result, the constant multiplier is no longer needed to convert base e (Euler's number) to other values because the base of GH CORDIC is adjustable. The proposed methodology is first validated using MATLAB with extensive vector matching. Then, example circuits with 16-bit fixed-point data are implemented under the TSMC 40-nm CMOS technology. Hardware experiment shows that at the highest frequency of the state of the art, the proposed methodology saves 27.98% area, 50.69% power consumption, and 6.67% latency when calculating logarithms; it saves 13.09% area, 40.05% power consumption, and 6.67% latency when computing exponentials. Both calculations do not compromise accuracy. Moreover, it can increase 13% maximum frequency and reduce up to 17.65% latency accordingly compared to the state of the art.
Yuanyong Luo, Yajun Ha, Zhongfeng Wang 0001, Hongbing Pan
IEEE Trans. Very Large Scale Integr. Syst.1
2019 Corrections to "Generalized Hyperbolic CORDIC and Its Logarithmic and Exponential Computation With Arbitrary Fixed Base"
abstract
In[1], the iterative formulas of generalized hyperbolic CORDIC, i.e.,(21), should read as follows:
Yuanyong Luo, Yajun Ha, Zhongfeng Wang 0001, Hongbing Pan
IEEE Trans. Very Large Scale Integr. Syst.1