Mingyu Shu

dblp:335/8920 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0000-9221-7669ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Efficient Hardware-Oriented Approximation of Mish Activation Function and Its Realization
Junting Liu, Mingyu Shu, Qiang Liu 0011
ISCAS2
2025 LHAM: Low-Cost and High-Accuracy Approximate Multiplier for FPGA-Based Computing
abstract
The rapid development of artificial intelligence raises higher demands on the performance of intelligent devices, requiring new computing units with lower resource usage and power consumption. Approximate Multipliers (AMs) meet this need by reducing resource and power consumption at the cost of computational accuracy and are widely used in fields like image processing and deep neural networks. In this article, we present a Low-Cost and High-Accuracy Approximate Multiplier (LHAM) design methodology targeting Field Programmable Gate Arrays (FPGAs). The expressions of carry propagation and carry generation for FPGA-based Carry-Look-Ahead Adders (CLA) are optimized, which effectively reducing the errors associated with discarding carry generation information. Using these expressions together with a logic fusion based approximation strategy, we design both accurate and approximate adders. These adders can be selectively configured during the partial product accumulation stage of the multiplier, allowing a tunable tradeoff between computational accuracy and hardware resource utilization. Finally, we model the AM design space as a 0-1 Knapsack problem to efficiently generate optimized designs under varying accuracy and area requirements. Experimental results show that, compared to the Xilinx accurate multiplier IP core, the proposed LHAM reduces LUT usage, delay, and power consumption by 50.7%, 21.7%, 20.6% for \(8\times 8\) multiplication, respectively. Compared to existing AMs, LHAM uses the fewest LUTs and achieves the best tradeoff between accuracy and area. The proposed LHAM is also applied to two applications to demonstrate its efficiency and effectiveness.
Mingyu Shu, Qiang Liu 0011
ACM Trans. Reconfigurable Technol. Syst.1
2024 PBN: Progressive Batch Normalization for DNN Training on Edge Device
abstract
Batch normalization (BN) plays a critical role in training deep neural networks (DNNs) on energy-limited edge devices since it accelerates the convergence of DNN training. However, the statistical operations and data dependencies within BN introduce challenges to efficient BN hardware design, such as complex computation and repeated data accesses. This paper presents PBN, a progressive batch normalization approach that can decouple the statistical calculation and normalization process within BN to address the above challenges. PBN exhibits considerable accuracy and convergence speed when evaluated using classical DNN models and datasets. Furthermore, a PBN hardware module that supports both forward propagation and backward propagation in DNN training is designed and implemented on Xilinx ZCU102 field-programmable gate array (FPGA). Experimental results indicate that PBN reduces external memory access (EMA) by an average of 80%, while achieving a 2.85× speedup compared with conventional BN.
Yingchang Mao, Mingyu Shu, Qiang Liu 0011
ISCAS2
2024 A Data-Distribution Aware Approximate Multiplier Design Based on FPGA
abstract
The approximate multiplier (AM) serves as a computing unit that saves hardware resources and power consumption at the expense of computational accuracy. This paper proposes a data-distribution aware approximate multiplier (DDAM) design for FPGAs. We build a numerical optimization model for automatic design space exploration of DDAM. Furthermore, we propose a weight-based iterative algorithm (WIA) to accelerate the solution of the optimization model. Experimental results demonstrate that WIA significantly reduces DDAM design exploration time to approximately 0.05% of a generic search method. The generated DDAM reduces the average error by about 43% compared to approximate multipliers with uniform data distribution. Furthermore, in an 8 × 8 multiplication scenario, DDAM reduces LUT utilization by 48% compared to the Xilinx’s accurate multiplier IP core. Compared to existing FPGA-based AMs, DDAM achieves the best balance between accuracy and area.
Mingyu Shu, Yingchang Mao, Qiang Liu 0011
ISCAS1
2022 LCAM: Low-Cost Approximate Multiplier Design on FPGA
abstract
Approximate multiplier is a computing unit, which reduces resource and power by sacrificing computational accuracy, and is widely used in fields such as image processing and deep neural networks. In this paper, a low-cost$\mathbf{8}\times \mathbf{8}$unsigned approximate multiplier is proposed by considering FPGA architectural features. A stage-aware most significant bits (MSBs) selection scheme is designed for error recovery to trade off accuracy and resource usage. The proposed multiplier saves up to 19.7% LUT utilization while the accuracy only decreases 4%, compared to the accurate Xilinx multiplier IP.
Mingyu Shu, Qiang Liu 0011
FPT1