Qianming Yang

dblp:35/6657 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-8796-6187ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Revisiting Global Value Prediction: A Resurgent Complement to Local Predictors
Ling Yang 0008, Libo Huang 0002, Bingcai Sui, Sheng Ma, Yongwen Wang, Li Shen 0007, Qianming Yang, Songwen Pei
ISCA10
2024 A Low-Cost Floating-Point Dot-Product-Dual-Accumulate Architecture for HPC-Enabled AI
abstract
The dot-product$\sum _{i=1}^{N} A_{i}\times B_{i}$is one of the most frequently used operations for a wide variety of high-performance computing (HPC) and artificial intelligence (AI) applications. However, for large-scale algorithms, such as acrshort GEMM and acrshort FFT, independent additions are necessary to accumulate the results of length-limited dot-product in order to form the final result, thus increasing latency and overhead. Hence, we proposed a dot-product-dual-accumulate (DPDAC) architecture capable of performing$\left({\sum _{i=1}^{N=1,2,4} A_{i}\times B_{i} + \sum _{j=1}^{M=1,2} C_{j}}\right)$on a wide range of formats. The proposed architecture supports both single-path and dual-path execution. The single path is designed for performing acrshort DP acrshort FMA or DPDAC of lower formats, while dual-path supports parallel operations for single-precision (SP) addition and 2-term SP or acrshort TF32 dot-product or 4-term acrshort HP or BF16 dot-product. Moreover, numerical precision conversion is also supported by the proposed architecture, allowing for the conversion of numbers to higher or lower formats. The proposed DPDAC has been demonstrated to significantly reduce the overhead in comparison to discrete designs that utilize multiple single-mode acrshort FP units to achieve the same functionalities. Furthermore, when compared to the state-of-the-art multiple-precision designs, the proposed architecture has been shown to support a wide range of formats and a greater variety of operations with lower costs.
Hongbing Tan, Libo Huang 0002, Hui Guo 0004, Qianming Yang, Li Shen 0007, Gang Chen 0023, Liquan Xiao, Nong Xiao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Low-Cost Multiple-Precision Multiplication Unit Design For Deep Learning
abstract
Low-precision formats have been proposed and applied to deep learning algorithms to speed up training and inference. This paper proposes a novel multiple-precision multiplication unit(MU) for deep learning. The proposed MU supports four types of precision for floating-point(FP) numbers-FP8-E4M3, FP8-E5M2, FP16, FP32-and 8-bit fixed-point(FIX) numbers. The MU can execute four parallel FP8 and eight parallel FIX8 multiplications simultaneously in one cycle, or four parallel FP16 multiplications fully pipelined with a latency of one, or one FP32 multiplication with a latency of one cycle. The simultaneous execution of FIX8 and FP8 can meet the requirements of the specific deep learning algorithms. Thanks to the low-precision-combination(LPC) and vectorization design method, multiplication in any precision can get 100% utilization of the multiplier resources, and the MU can adopt a lower clock delay to achieve better performance in all data types. Compared with the existing multiple-precision units designed for deep learning, this MU can support more types of low-precision formats by lower area overhead; and exhibits higher throughput at FIX8 with at least 8× improvement.
Libo Huang 0002, Hongbing Tan, Ling Yang 0008, Qianming Yang
ACM Great Lakes Symposium on VLSI6
2022 Efficient Multiple-Precision and Mixed-Precision Floating-Point Fused Multiply-Accumulate Unit for HPC and AI Applications
Hongbing Tan, Run Yan, Ling Yang 0008, Libo Huang 0002, Liquan Xiao, Qianming Yang
ICA3PP6
2017 Optimizing OpenCL Implementation of Deep Convolutional Neural Network on FPGA
Yuran Qiao, Junzhong Shen, Dafei Huang, Qianming Yang, Mei Wen, Chunyuan Zhang
NPC4
2017 FPGA-accelerated deep convolutional neural networks for high throughput and energy efficiency
abstract
Summary Recent breakthroughs in the deep convolutional neural networks (CNNs) have led to great improvements in the accuracy of both vision and auditory systems. Characterized by their deep structures and large numbers of parameters, deep CNNs challenge the computational performance of today. Hardware specialization in the form of field‐programmable gate array offers a promising path towards major leaps in computational performance while achieving high‐energy efficiency. In this paper, we focus on accelerating deep CNNs using the Xilinx Zynq‐zq7045 FPGA SoC. As most of the computational workload can be converted to matrix multiplications, we adopt a matrix multiplier‐based accelerator architecture. Dedicated units are designed to eliminate the conversion overhead. We also design a customized memory system according to the memory access pattern of CNNs. To make the accelerator easily usable by application developers, our accelerator supports Caffe, which is a widely used software framework of deep CNN. Different CNN models can be adopted by our accelerator, with good performance portability. The experimental results show that for a typical application of CNN, image classification, an average throughout of 77.8 GFLOPS is achieved, while the energy efficiency is 4.7× better than an Nvidia K20 GPGPU. © 2016 The Authors. Concurrency and Computation: Practice and Experience Published by John Wiley & Sons Ltd
Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen, Chunyuan Zhang
Concurr. Comput. Pract. Exp.4
2015 Unified Virtual Memory Support for Deep CNN Accelerator on SoC FPGA
Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen
ICA3PP (1)4
2013 Accelerating thread-intensive and explicit memory management programs with dynamic partial reconfiguration
Qianming Yang, Mei Wen, Nan Wu 0003, Chunyuan Zhang
J. Supercomput.1
2012 The masala machine: accelerating thread-intensive and explicit memory management programs with dynamically reconfigurable FPGAs (abstract only)
abstract
A uniform FPGA-based architecture, an efficient programming model and a simple mapping method are paramount for PPGA technology to be more widely accepted. This paper presents MASALA, a dynamically reconfigurable FPGA-based accelerator specifically for parallel programs written in thread-intensive and explicit memory management (TEMM) programming models. The system uses TEMM programming model to parallelize the demanding application, including decomposing the application into separate thread blocks, decoupling compute and data load/store etc. Hardware engines are included into the MASALA by using partial dynamic reconfigure modules, each of which encapsulates Thread Process Engine implementing the thread functionality in hardware. A data dispatching scheme is also included in MASALA to enable the explicit communication among multiple memory hierarchies such as between inter-hardware engines, the host processor and hardware engines. At last, the paper illustrates a Multi-FPGA prototype system of the presented architecture: MASALA-SX. A large synthetic aperture radar (SAR) image formatting experiment shows that the MASALA architecture facilitates the construction of a TEMM program accelerator by providing it with greater performance and less power consumption than current CPU platforms, but without sacrificing programmability, flexibility and scalability.
Mei Wen, Nan Wu 0003, Qianming Yang, Chunyuan Zhang
FPGA3
2010 Software Managed Instruction Scratchpad Memory Optimization in Stream Architecture Based on Hot Code Analysis of Kernels
abstract
Stream processors, such as Imagine, GPGPUs, FT64 and MASA, typically uses software managed scratchpad instruction memory which improves performance and significantly reduces energy consumption. In this paper, we build a kernel-storage model to analyze the hot spot of kernels in stream programs. Based on the analysis, we define Kernel Hot Code and prove that scratchpad instruction memory should focus on the access efficiency of it. A methodology for finding Kernel Hot Code in the kernels of different structures is presented as well. In accordance with this method, we develop HOIS for Stream Architecture, which adopts a software managed scratchpad memory to store Kernel Hot Code, and uses a small hardware managed victim cache to store the Kernel Cool Code. HOIS is evaluated by measuring the performance of six applications on the MASA_S simulation platform. The results show that HOIS can achieve high efficiency in predictable applications with little performance loss.
Yi He 0008, Ju Ren 0002, Mei Wen, Qianming Yang, Nan Wu 0003, Chunyuan Zhang
DSD4
2007 FT64: Scientific Computing with Streams
Mei Wen, Nan Wu 0003, Chunyuan Zhang, Qianming Yang, Changqing Xun
HiPC5