VLDB 2026 Research / reviewers in the wild / expert
Yu Wang 0161
dblp:02/5889-161
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-5561-4435ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Universal Methodology of Complex Number Computation for Low-Complexity and High-Speed ImplementationabstractIn complex-valued neural network (CVNN) applications, complex number calculations require high performance rather than high precision. However, most previous studies focused on high-precision approaches, which have low speed and high hardware costs. This paper proposes a universal methodology of complex number computation for low-complexity and high-speed implementation. The proposed methodology is based on the piecewise linear (PWL) method and can be used for different types of complex number computations. Considering that multiplication operations consume considerable resources, multiplication, fused square-add (FSA) and fused multiply-add (FMA) operations are the focus of optimization. The partial products of the square operation are reduced by folding and merging techniques because of their symmetry in the FSA operation. The partial products of the multiplication and FMA operations are reduced via Booth encoding. In addition, the partial products are further reduced by the proposed step-by-step truncation method. The proposed segmenter, which simulates the hardware implementation, automatically divides the nonlinear functions in the complex number computations into the smallest number of segments according to the required precision. The results show that the proposed approach improves performance and reduces hardware costs compared with the state-of-the-art methods for complex number calculations involving square roots, reciprocals and logarithms. Yu Wang 0161, Youlong Wu, Fei Lyu 0002, Yuanyong Luo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | An Optimized Architecture for Computing the Square Root of Complex NumbersabstractIn this paper, we propose an optimized architecture for computing the square root of complex numbers based on the piecewise linear (PWL) method. In the proposed design, the square-add and multiply-add operations are the focus of optimization. The symmetry of the partial products of the square operation is used to reduce the number of partial products. In addition, the least significant bits (LSBs) of the square-add and multiply-add are truncated to reduce the bits of partial products. According to the optimization of the hardware circuit, the simulation of the circuit in the segmentor is modified by introducing the simulation of the truncated fused square-add operation and truncated fused multiply-add operation. Experimental results show that the proposed optimization architecture has superiority in area, delay and power when compared with state-of-the-art designs. Yu Wang 0161, Fei Lyu 0002, Yuanyong Luo |
ISCAS | 1 |
| 2022 | Reconfigurable Multifunction Computing Unit Using an Universal Piecewise Linear MethodabstractComputing units for nonlinear complex functions are indispensable in deep neural network training processors. However, the existing computing units for nonlinear complex functions have low utilization efficiency and poor agreement and precision. In this article, we propose a multifunction computing unit for training deep neural networks by reusing computing resources based on a piecewise linear (PWL) method to improve computing density. Based on the state-of-the-art segmentor of PWL method, multiple nonlinear functions are divided into the fewest segments with the same bit width of computation. In hardware implementation, the reconfigurable technique is implemented on multiple functions while reusing computing resources including the multiplier and adder. The application-specific integrated circuit (ASIC) implementation results reveal that the architecture with reuse reduces the area by 44.50% and the power by 43.71% at the same frequency, when compared with the architecture without reuse. Fei Lyu 0002, Wenxiu Wang, Yuanyong Luo, Yu Wang 0161 |
ISCAS | 6 |
| 2022 | ML-PLAC: Multiplierless Piecewise Linear Approximation for Nonlinear Function EvaluationabstractIn this article, we propose a multiplierless piecewise linear (PWL) approximation computation (ML-PLAC) method for nonlinear unary functions. ML-PLAC seeks the minimum number of segments with the predefined fractional bit width and number of adders to satisfy the restriction on the maximum absolute error (MAE). Compared with the previous universal PWL approximation method, multiplication operations in the segmentor are replaced by a simulation of shift-and-add operations by reducing the fractional bit width of the slope of linear functions. Various numbers of segments are obtained by different predefined numbers of adders to balance the two numbers. In addition, adders are used to replace the multiplier in hardware architecture. The synthesized results prove that ML-PLAC has increased performance without any compromises. Compared with state-of-the-art methods, ML-PLAC saves 58.82% area, 38.16% delay, 60.31% power and 6.10%MAEwhen computing logarithmic functions; 51.51% area, 46.49% delay, and 46.95% power while maintaining the comparableMAEwhen computing antilogarithmic functions; 55.21% area, 25% delay, 61.21% power, and 37.30%MAEwhen computing hyperbolic tangent functions; 82.47% area, 60% delay, 77.51% power and 12.74%MAEwhen computing sigmoid functions; and 46.43% area, 31.43% delay, and 61.16% power while maintaining the sameMAEwhen computing softsign functions. Fei Lyu 0002, Zhelong Mao, Yanxu Wang, Yu Wang 0161, Yuanyong Luo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | High-Throughput Low-Latency Pipelined Divider for Single-Precision Floating-Point NumbersabstractIn this brief, we propose a fully pipelined divider for single-precision floating-point numbers based on a universal piecewise linear (PWL) approximation method and a modified Goldschmidt algorithm. The state-of-the-art universal PWL method uses a suitable number of segments and fractional bit widths to meet the requirement of the predefined maximum absolute error. Small multipliers are employed in the modified Goldschmidt algorithm. In the hardware implementation, the multipliers are optimized with the radix-4 and radix-8 booth encoding methods to reduce the number of partial products. In addition, the sum of the partial products and other data are calculated by a compressor and an adder to shorten the critical path. Synthesized results show that the maximum achievable frequency of our design is better than those of the existing methods. In addition, our design shows overwhelming superiority in terms of latency and throughput compared with existing methods. Fei Lyu 0002, Yanxu Wang, Yuanyong Luo, Yu Wang 0161 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2021 | Ultralow-Latency VLSI Architecture Based on a Linear Approximation Method for Computing Nth Roots of Floating-Point NumbersabstractState-of-the-art approaches that perform root computations based on the COordinate Rotation Digital Computer (CORDIC) algorithm suffer from high latency in performing multiple iterations. Therefore, root computations based on the CORDIC algorithm cannot meet the strict latency requirements of some applications. In this paper, we propose a methodology for performing Nth root computations on floating-point numbers based on the piecewise linear (PWL) approximation method. The proposed method divides an Nth root computation into several subtasks approximated by the PWL algorithm. It determines the widest segments of the subtasks and the smallest fractional width needed to satisfy the predefined maximum relative error Max_Errr. Our design is coded in Verilog HDL and synthesized under TSMC 40 nm CMOS technology. The synthesized results show that our design can reach the highest frequency of 2.703 GHz with an area consumption of 2608.84 μ m2and a power consumption of 2.4476 mW. Compared with one stateof-the-art architecture, our design saves 91.60%, 89.84%, and 63.33% of the area, power, and latency @1.89GHz frequency, respectively, while reducing Max_Errrby 57.30%. In addition, it saves 94.52%, 92.68%, and 73.17% of the area, power, and delay @1.89GHz frequency, respectively, and reduces Max_Errrby 1.65% when compared with the other state-of-the-art design. Fei Lyu 0002, Xiaoqi Xu, Yu Wang 0161, Yuanyong Luo, Hongbing Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | PWL-Based Architecture for the Logarithmic Computation of Floating-Point NumbersabstractIn this brief, we propose a logarithmic converter for floating-point numbers based on the piecewise linear (PWL) approximation method. The proposed method is applicable to any customized floating-point format with a mantissa length of 16–23 bits and a maximum absolute error (MAE) larger than 10−6. The logarithmic function is automatically segmented into several maximal subsections by a software-based segmentation scheme with the restriction of a predefined MAE and a fractional word length for the computing units. Then, we make a tradeoff between the piecewise number and the fractional word length. Based on the results of the segmentor, our design is coded in the Verilog hardware description language. The synthesized results show that our design consumes less area, time, and power without compromising accuracy compared to existing techniques based on the COordinate Rotation Digital Computer (CORDIC) and PWL methods. Fei Lyu 0002, Zhelong Mao, Yu Wang 0161, Yuanyong Luo |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |