Yongming Tang

dblp:55/2826 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-2102-2041ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 8 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Fine-grained data integration for high throughput and bandwidth-efficient computation on FPGAs
Jiyuan Liu 0006, Baoping Wang, Yongming Tang, He Li 0008
Integr.3
2025 Scalable and Real-Time Power System Simulation Based on Heterogeneous CPU-FPGA Co-operation
abstract
With the increasing integration of renewable energy devices, modern power systems have become complex, exhibiting diverse circuit topologies. Existing FPGA-based accelerators are optimized for fast single-topology simulation but lack the flexibility to handle multiple topology simulations. This paper proposes a scalable heterogeneous CPU-FPGA system designed for efficient simulation of diverse power system topologies. The proposed system enables real-time simulation for variable-scale power systems, distinguishing itself from conventional simulators by leveraging flexible matrix decomposition and a topology-aware approach. Our design specifically achieves the minimal latency of 100ns for both the boost and three-phase voltage source converter (VSC) models, highlighting its superior performance across diverse topologies. Introduced by dynamic fixed-point quantization, our proposal reduces 30.8% LUTs, 26.8% FFs, 22.2% BRAMs and 43.8% DSPs versus a fixed-point implementation with the same computational accuracy.
Hangyu Yang, Jiyuan Liu 0006, Mingwang Xu, Yongming Tang, He Li 0008
ISCAS5
2023 MSBF-LSTM: Most-significant Bit-first LSTM Accelerators with Energy Efficiency Optimisations
abstract
Long short-term memory (LSTM) recurrent networks are frequently applied to sequence processing problems such as speech recognition and video classification. In this paper, we propose a novel LSTM network implementation based on most-significant bit-first (MSBF) arithmetic on FPGAs, called MSBF-LSTM, to improve the energy efficiency of LSTM inference engines. Furthermore, an LSTM model compression strategy with incremental network quantization is proposed to achieve high accuracy with low-precision weights. Hardware implementations conducted on the Xilinx UltraScale+ Zynq xczu9eg FPGA demonstrate that MSBF-LSTM achieves 1.51× better energy efficiency compared with the state-of-the-art FPGA-based LSTM designs.
Sige Bian, He Li 0008, Changjun Song, Yongming Tang
FCCM5
2023 MSDF-SGD: Most-Significant Digit-First Stochastic Gradient Descent for Arbitrary-Precision Training
abstract
Stochastic gradient descent has been a widely used machine learning algorithm, and interest in low-precision SGD is growing because it improves throughput and keeps efficient convergence. We propose MSDF-SGD, a novel approach allowing SGD to support arbitrary-precision training on FPGAs by employing most-significant digit-first arithmetic. MSDF-SGD is the first architecture that supports arbitrary-precision data, models, and intermediates at the same time. MSDF-SGD is evaluated via training linear classifiers on representative datasets. MSDF-SGD delivers a 1.6× speedup over state-of-the-art low-precision hardware implementations and converges up to 8.6× faster than cutting-edge implementations on CPUs. Finally, we provide a programming interface that permits building a custom arbitrary-precision training accelerator, making MSDF-SGD support more complicated, multi-layered and nonlinear models.
Changjun Song, Yongming Tang, Jiyuan Liu 0006, Sige Bian, Danni Deng, He Li 0008
FPL2
2023 Design Space Exploration for Efficient Quantum Most-Significant Digit-First Arithmetic
abstract
Quantum computing has been considered as an emerging approach in addressing problems which are not easily solvable using classical computers. In parallel to the physical implementation of quantum processors, quantum algorithms have been actively developed for real-life applications to show quantum advantages, many of which benefit from quantum arithmetic algorithms and their efficient implementations. As one of the most important operations, quantum addition has been adopted in Shor's algorithm and quantum linear algebra algorithms. Although various least-significant digit-first quantum adders have been introduced in previous work, interest in investigating the efficient implementation of most-significant digit-first addition is growing. In this work, we propose a novel design method for most-significant digit-first addition with several quantum circuit optimisations to reduce the number of quantum bits (i.e. qubits), quantum gates, and circuit depth. An open-source library of different arithmetic operators based on our proposed method is presented, where all circuits are implemented on IBM Qiskit SDK. Extensive experiments demonstrate that our proposed design, together with the optimisation techniques, reduces T-depth by up-to 4.0×, T-count by 3.5×, and qubit consumption by 1.2×.
He Li 0008, Hongxiang Fan, Yongming Tang
IEEE Trans. Computers4
2021 Real-Time Super-Resolution System of 4K-Video Based on Deep Learning
abstract
Video super-resolution (VSR) technology excels in reconstructing low-quality video, avoiding unpleasant blur effect caused by interpolation-based algorithms. However, vast computation complexity and memory occupation hampers the edge of deplorability and the runtime inference in real-life applications, especially for large-scale VSR task. This paper explores the possibility of real-time VSR system and designs an efficient and generic VSR network, termed EGVSR. The proposed EGVSR is based on spatio-temporal adversarial learning for temporal coherence. In order to pursue faster VSR processing ability up to 4K resolution, this paper tries to choose lightweight network structure and efficient upsampling method to reduce the computation required by EGVSR network under the guarantee of high visual quality. Besides, we implement the batch normalization computation fusion, convolutional acceleration algorithm and other neural network acceleration techniques on the actual hardware platform to optimize the inference process of EGVSR network. Finally, our EGVSR achieves the real-time processing capacity of [email protected]. Compared with TecoGAN, the most advanced VSR network at present, we achieve 85.04% reduction of computation density and 7.92× performance speedups. In terms of visual quality, the proposed EGVSR tops the list of most metrics (such as LPIPS, tOF, tLP, etc.) on the public test dataset Vid4 and surpasses other state-of-the-art methods in overall performance score.
Yanpeng Cao, Changjun Song, Yongming Tang, He Li 0008
ASAP4
2021 A General Video Processing Framework on Edge Computing FPGAs
abstract
Digital video processing needs high bandwidth transmission from source to host, which poses a massive challenge to existing technologies. As an appropriate solution, edge computing can provide immediate process with low latency, high transmission bandwidth and memory usage. In this paper, we propose a general video processing framework on edge computing FPGAs (GVPF-E), consisting of in-out buffers and an update-feedback mechanism. For general video processing algorithms, GVPF-E extracts multiple video frames' inter-correlation and updates calculation results into output feature buffers, so as to we can optimize the video processing quality via update-feedback mechanism. Our illustrative hardware implementations on a CNN-based video filtering algorithm achieve an up-to 3.87TMACS performance under 8-bit quantization. We also obtain 2.45GB/s bandwidth and 94.4% peak throughput utilization under Xilinx XCZU15EG embedded computing FPGAs.
Feng Yu 0006, He Li 0008, Rongshi Dai, Yongming Tang
FCCM4
2021 ASIC Design Principle Course with Combination of Online-MOOC and Offline-Inexpensive FPGA Board
abstract
ASIC Design Principle (ASICDP) is a compulsory course for undergraduate majors in microelectronics and integrated circuits, and the focus of this paper is the teaching methods of online theoretical teaching and offline experimental teaching of this course. As is well known, in order to prevent and control COVID-19, the use of online platforms to carry out online teaching has attracted worldwide attention. In this paper, the teaching strategy "Online-MOOC + Offline Inexpensive FPGA Board" in ASICDP in the Spring 2020 semester is demonstrated, in where MOOC means Massive Open Online Course. The theoretical teaching content of ASICDP is entirely replicated from Hardware Acceleration Design Methodology (HADM) released by the present authors on "China University MOOC," the largest MOOC platform in China. Meanwhile, with the support of the "Xilinx & Ministry of Education University-Industry Collaborative Education Program," an FPGA development board called the "Spartan Edge Accelerator Board" (SEA Board) designed by the authors was used in the experimental teaching of the ASICDP. This method can be used to establish the linkage between online courses and offline experiments, and cultivate students' practical VLSI design and FPGA prototype verification skills. It is believed that for educators that want to improve courses related to ASIC design and FPGA prototype verification, Online-MOOC + Offline-Inexpensive FPGA Board is an effective method with lower cost that is easily promotable and replicated.
Zhixiong Di, Yongming Tang, Jiahua Lu, Zhaoyang Lv
ACM Great Lakes Symposium on VLSI2
2020 Explore Efficient LUT-based Architecture for Quantized Convolutional Neural Networks on FPGA
abstract
The vast computations of the convolutional neural network have limited the speed of the forward inference running in hardware. In recent years, network quantization technique has made it possible to quantize network into low bit-wide and retain the original performance simultaneously, while the complexity of the quantized network is still considerable. FPGA is a highly parallelized platform, which contains a mass of configurable logic resources. We study on the feasibility of implementing convolution calculation based on pure LUTs, introduce the shift multipliers and addition trees, and propose an efficient architecture for QNN on FPGA. With the optimization of Winograd algorithm for QNN, we demonstrate that our scheme significantly reduces the number of multipliers and saves the usage of LUT resources by $2.25 \times $ at least without using DSP resources. As a result, our LUT-based architecture for QNN shortens the latency up to $19.3 \times $ and represents more effective performance compared to other methods.
Yanpeng Cao, Yongming Tang
FCCM3
2020 Realization of Quantized Neural Network for Super-resolution on PYNQ
abstract
Vision tasks usually require vast amount of computation and memory resources, which create barriers to edge computing applications. Quantized neural network can provide memory saving, scalability and energy efficiency, while the accuracies of results may decrease. In this paper, we adjust the data-width of feature maps, weights and temporary variables in SRCNN to achieve a trade-off between precision and accuracy. Also, we design a dedicated convolutional acceleration for data stream under the heterogeneous CPU-FPGA platform: PYNQ, including the changed data streaming order and im2col for convolution. Results show when data-width was set to 12-bit, quantization had almost no effect on visual perception of superresolved images. The acceleration of quantized convolution on FPGA can achieve a speed up ratio of 120x at 250MHz, compared with ARM CPU.
Feng Yu 0006, Yanpeng Cao, Yongming Tang
FCCM3
2014 Cross-cultural active learning: Qualitative results from Americans teaching in China
abstract
What are the experiences of Chinese students taking engineering courses taught in English by American professors using active learning techniques? Case studies of three such courses are explored in this work. Specifically, two American professors taught courses in English to about 100 students whose native language was Chinese. These courses included a required Electronics course for electrical engineering sophomores, a seminar on Medical Device New Product Development within a required Biomedical Instrumentation course for juniors in biomedical engineering, and a lecture series on New Product Development for first year honors students. All courses included homework teams and active learning techniques in the classroom. Focus groups were conducted near the end of the courses to examine students' experiences. Results and analyses of this qualitative data are presented here. Themes which emerged include the importance of teamwork, challenges with and comfort gained using English, conflicting attitudes toward teaching style, and exposure to practical applications.
Susan M. Lord, Victor W. Chang, Yinghui Kuang, Yongming Tang
EDUCON5
2013 Cross-cultural active learning: Results from Americans teaching in China
abstract
What are the experiences of Chinese students taking engineering courses taught in English by American professors using active learning techniques? Case studies of three such courses are explored in this work. Specifically, two American professors taught courses in English to about 90 students whose native language was Chinese. These courses included a required Electronics course for electrical engineering sophomores, a seminar on Medical Device Product Development within a required Biomedical Instrumentation course for juniors in biomedical engineering, and a lecture series on New Product Development for first year honors students. All courses included homework teams and active learning techniques in the classroom. Surveys were given to students throughout the courses. Statistical analyses of these surveys show that the Chinese students valued the homework team experience and believed that it helped them learn although some variation is seen among the courses. Peer evaluations of teamwork also showed that students responded well to homework teams. Despite considerable variation in their comfort with speaking English, the students did well in these courses taught in English with active learning techniques.
Susan M. Lord, Rick T. Olson, Yinghui Kuang, Victor W. Chang, Yongming Tang
EDUCON5