Yingchang Mao

dblp:369/4872 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-2731-3571ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 ALN: Approximate Layer Normalization for Transformer Training on Edge Device
Yingchang Mao, Qiang Liu 0011
IEEE Trans. Computers1
2025 SDTA: An Efficient Sparse DNN Training Accelerator with Data Hierarchical Pre-fetching and Dynamic Scheduling
abstract
Recently, training deep neural networks (DNNs) on edge devices has attracted much attention due to its strong adaptability and avoidance of private data transmission. However, limited computational, storage, and energy resources pose significant challenges for edge devices. The structural and computational redundancies in DNNs create opportunities for sparse training through model pruning and zero-computation skipping. Although feasible, the sparse training accelerator design encounters common issues, such as redundant data duplication and unbalanced workloads, caused by irregular sparsity. To address these issues, this paper proposes a sparse DNN training accelerator, SDTA, together with a hierarchical pre-fetching buffer and a dynamic scheduler to achieve high design efficiency. The SDTA is deployed on the FPGA XCVU3P platform. Compared to the prior FPGA-based accelerators and the GPU, SDTA improves the energy efficiency by up to 2.29×, the storage utilization efficiency by up to 7.37×, and the computational efficiency by up to 1.9×. Compared to the dense accelerator, it achieves a speedup of up to 5.88×, while ensuring model accuracy.
Mengting Wang, Yuntao Han, Yingchang Mao, Peng Shao, Zhengyan Liu, Qiang Liu 0011
ISCAS3
2024 MSCA: A Multi-Grained Sparse Convolution Accelerator for DNN Training
abstract
Training deep neural networks (DNNs) on edge devices is appealing for its adaptability and privacy benefits, but it faces challenges due to the limited resources and energy available on edge devices. In this paper, we propose MSCA, a Multigrained Sparsity Convolution Accelerator. MSCA exploits both coarse-grained and fine-grained sparsity during the DNN training phases through two types of well-designed units. Experimental results show that MSCA implemented on FPGA achieves 218.03 GOPS throughput, 39.8 GOPS/W energy efficiency, and 4.0-6.2x speedup over dense accelerators for training VGG-8 and ResNet-10 on the CIFAR-10 and SVHN datasets.
Yingchang Mao, Qiang Liu 0011, Ray C. C. Cheung
ASAP1
2024 PBN: Progressive Batch Normalization for DNN Training on Edge Device
abstract
Batch normalization (BN) plays a critical role in training deep neural networks (DNNs) on energy-limited edge devices since it accelerates the convergence of DNN training. However, the statistical operations and data dependencies within BN introduce challenges to efficient BN hardware design, such as complex computation and repeated data accesses. This paper presents PBN, a progressive batch normalization approach that can decouple the statistical calculation and normalization process within BN to address the above challenges. PBN exhibits considerable accuracy and convergence speed when evaluated using classical DNN models and datasets. Furthermore, a PBN hardware module that supports both forward propagation and backward propagation in DNN training is designed and implemented on Xilinx ZCU102 field-programmable gate array (FPGA). Experimental results indicate that PBN reduces external memory access (EMA) by an average of 80%, while achieving a 2.85× speedup compared with conventional BN.
Yingchang Mao, Mingyu Shu, Qiang Liu 0011
ISCAS1
2024 A Data-Distribution Aware Approximate Multiplier Design Based on FPGA
abstract
The approximate multiplier (AM) serves as a computing unit that saves hardware resources and power consumption at the expense of computational accuracy. This paper proposes a data-distribution aware approximate multiplier (DDAM) design for FPGAs. We build a numerical optimization model for automatic design space exploration of DDAM. Furthermore, we propose a weight-based iterative algorithm (WIA) to accelerate the solution of the optimization model. Experimental results demonstrate that WIA significantly reduces DDAM design exploration time to approximately 0.05% of a generic search method. The generated DDAM reduces the average error by about 43% compared to approximate multipliers with uniform data distribution. Furthermore, in an 8 × 8 multiplication scenario, DDAM reduces LUT utilization by 48% compared to the Xilinx’s accurate multiplier IP core. Compared to existing FPGA-based AMs, DDAM achieves the best balance between accuracy and area.
Mingyu Shu, Yingchang Mao, Qiang Liu 0011
ISCAS2