Yongsheng Yin

dblp:277/4601 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 A calibration technology for SAR ADC with PGA based on poly resistor linearization and sampling capacitance bit-weight mismatch
Wei Sheng, Jiashen Li, Yan Xue, Mingyuan Ye, Yongsheng Yin
Integr.6
2026 Genetic algorithm-optimized fuzzy controller for the calibration of pipelined ADCs
Luotian Wu, Honghui Deng, Jiashen Li, Muqi Li, Yongsheng Yin
Integr.6
2026 A Neural Network-Based ADC Calibration Framework via Quantization Code Reconstruction
abstract
This article proposes a neural network (NN)-based calibration framework via quantization code reconstruction to address the critical limitation of multidimensional NNs (MDNNs) in analog-to-digital converter (ADC) calibration, where performance degrades drastically under varying input frequencies or sampling rates. Conventional MDNN methods suffer from distribution shifts caused by dynamic oversampling ratio (OSR) variations, necessitating repeated retraining. Our innovation lies in a data reconstruction mechanism: for lower OSR scenarios, a third-order cascaded integrator-comb (CIC) filter inserts pseudo-data points between quantization code to match the NN’s input dimension. For higher OSR scenarios, cyclic downsampling decomposes the original sequence into parallel subsequences processed independently by the same network. For multitone signals, a reconstruction–calibration–summation flow hierarchically handles spectral components. Implemented on an FPGA and covalidated with a commercial 14-bit/1 GSPS pipeline ADC, the framework demonstrates robust performance without retraining. Experimental results show: bandwidth expansion from 16.3% to 84.7% of the Nyquist bandwidth. SFDR improvement of 13.8–27.7 dB and ENOB gain of 2.3–4.6 bits across 10.3–160.3 MHz inputs. This work enables consistent high-precision ADC calibration in dynamic signal environments, facilitating deployment in complex application scenarios.
Jiashen Li, Honghui Deng, Muqi Li, Luotian Wu, Xiaoting Lu, Yongsheng Yin
IEEE Trans. Very Large Scale Integr. Syst.8
2026 ParaPM: Efficient Hardware Accelerator for Postquantum Signature With High-Performance Polynomial Multiplier
abstract
During the Institute of Standards and Technology (NIST) postquantum cryptography standardization process, the lattice-based Dilithium scheme was selected as one of the three third-round finalists for digital signature algorithms. Although numerous hardware implementations of Dilithium have been proposed, there remains substantial room for performance optimization, particularly in terms of computational speed. In this work, we target the two most time-consuming operations, namely the coefficient generation and polynomial multiplication. We propose a high-speed hardware architecture that fully exploits hardware parallelism. Our design introduces a fully pipelined radix-2 multipath delay commutator (R2MDC) structure supporting both NTT and inverse NTT (INTT) modes, a throughput-matched polynomial multiplier enabling on-the-fly pointwise multiplication (PWM) without buffering, and a low-resource Keccak. These optimizations collectively eliminate the throughput mismatch bottleneck and maximize hardware utilization through task-level parallelism. We implement coefficient generation, signing, and verification for three security levels on the A7 and Z7 platforms. Experimental results demonstrate that our current design achieves the optimal area-time product (ATP) across all three security levels.
Gaoming Du, Yongsheng Yin, Duoli Zhang, Zhenmin Li
IEEE Trans. Very Large Scale Integr. Syst.4
2025 Digital background calibration algorithm for pipelined ADC based on time-delay neural network with genetic algorithm feature selection
Yongsheng Yin, Jiashen Li, Honghui Deng, Hongmei Chen 0005, Luotian Wu, Muqi Li
Integr.1
2022 A BNN Accelerator Based on Edge-skip-calculation Strategy and Consolidation Compressed Tree
abstract
Binarized neural networks (BNNs) and batch normalization (BN) have already become typical techniques in artificial intelligence today. Unfortunately, the massive accumulation and multiplication in BNN models bring challenges to field-programmable gate array (FPGA) implementations, because complex arithmetics in BN consume too much computing resources. To relax FPGA resource limitations and speed up the computing process, we propose a BNN accelerator architecture based on consolidation compressed tree scheme by combining both XNOR and accumulation operation of the low bit into a systematic one. During the compression process, we adopt 0-padding (not ±1) to achieve no-accuracy-loss from software modeling to hardware implementation. Moreover, we introduce shift-addition-BN free binarization technique to shorten the delay path and optimize on-chip storage. To sum up, we drastically cut down the hardware consumption while maintaining great speed performance with the same model complexity as the previous design. We evaluate our accelerator on MNIST and CIFAR-10 dataset and implement the whole system on the ARTIX-7 100T FPGA with speed performance of 2052.65 GOP/s and area efficiency of 70.15 GOPS/KLUT.
Gaoming Du, Bangyi Chen, Zhenmin Li, Zhenxing Tu, Shenya Wang, Qinghao Zhao, Yongsheng Yin
ACM Trans. Reconfigurable Technol. Syst.8
2020 A split-based fully digital feedforward background calibration technique for timing mismatch in TIADC
Hongmei Chen 0005, Yongsheng Yin, Linhao Gan, Honghui Deng
Integr.2
2020 Application of 3D laser scanning technology for image data processing in the protection of ancient building sites through deep learning
Yongsheng Yin, Juan Antonio
Image Vis. Comput.1
2019 Efficient Softmax Hardware Architecture for Deep Neural Networks
abstract
Deep neural network (DNN) has become a pivotal machine learning and object recognition technology in the big data era. The softmax layer is one of the key component layers for completing multi-classification tasks. However, the softmax layer contains complex exponential and division operations, resulting in low accuracy and long critical paths in hardware accelerator design. In order to solve the above issues, we present a softmax hardware architecture with proper accuracy, good trade-off and strong expansibility. We summarize the classification rules of neural network and balance the calculation accuracy between resource consumption. On this basis, we proposed an exponential calculation unit based on the group lookup table, and improve a natural logarithmic calculation unit based on the Maclaurin series and the data preprocessing scheme matching them. The experimental results show that the softmax hardware architecture proposed in this paper can achieve the calculation accuracy of 3 decimal fraction and the classification accuracy of $99.01%$. Theoretically, it can accomplish the classification task of infinite categories.
Gaoming Du, Zhenmin Li, Duoli Zhang, Yongsheng Yin
ACM Great Lakes Symposium on VLSI5
2019 NR-MPA: Non-Recovery Compression Based Multi-Path Packet-Connected-Circuit Architecture of Convolution Neural Networks Accelerator
abstract
Convolution Neural Networks (CNNs) involve massive data to be calculated and stored. To meet the challenges above, parallel hardware accelerators consisting of hundreds of Processing Elements (PEs) arranged as a many-core systemon-chip, connected by a Network-on-Chip (NoC) are proposed, which achieve high throughput exploiting parallel PE array. However, most of existing accelerators focus on only one aspect, such as compute structure of PE and data movement overhead above NoC, which causes the throughout, area and latency of the accelerator not fully optimized. In this paper, we propose an efficient general purpose CNN accelerator including both compute based on Non-Recovery Compression (NRC) method and data movement by novel Multi-Paths Packet Connection Circuit (MP-PCC). NRC can save computation time due to zero multiplier through shift decoding in PE and improve power efficiency by saving a large number of data transmission. MPPCC, evolved from Packet Connection Circuit, supports single and multicast transmission modes at the same time, and changes the multicast (X, Y) routing algorithm to multicast Y algorithm to improve the transmission efficiency. The proposed architecture which was implemented on Xilinx FPGA achieves 17.7x faster computation speed and 2.2x fewer memory accesses compared with the state-of-the-art method.
Gaoming Du, Zhenwen Yang, Zhenmin Li, Duoli Zhang, Yongsheng Yin, Zhonghai Lu
ICCD5
2017 All-digital background calibration technique for timing mismatch of time-interleaved ADCs
Hongmei Chen 0005, Yunsheng Pan, Yongsheng Yin, Fujiang Lin
Integr.3
2007 On the Implementation of Virtual Array Using Configuration Plane
Yongsheng Yin, Li Li 0003, Minglun Gao, Gaoming Du, Yu-Kun Song
APPT1