EDBT 2026 Demo / reviewers in the wild / expert
Qianming Yang
dblp:35/6657
· DBLP profile ↗
11ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-8796-6187ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Global Value Prediction: A Resurgent Complement to Local Predictors
Ling Yang 0008, Libo Huang 0002, Bingcai Sui, Sheng Ma, Yongwen Wang, Li Shen 0007, Qianming Yang, Songwen Pei |
ISCA | 10 |
| 2024 | A Low-Cost Floating-Point Dot-Product-Dual-Accumulate Architecture for HPC-Enabled AIabstractThe dot-product$\sum _{i=1}^{N} A_{i}\times B_{i}$is one of the most frequently used operations for a wide variety of high-performance computing (HPC) and artificial intelligence (AI) applications. However, for large-scale algorithms, such as acrshort GEMM and acrshort FFT, independent additions are necessary to accumulate the results of length-limited dot-product in order to form the final result, thus increasing latency and overhead. Hence, we proposed a dot-product-dual-accumulate (DPDAC) architecture capable of performing$\left({\sum _{i=1}^{N=1,2,4} A_{i}\times B_{i} + \sum _{j=1}^{M=1,2} C_{j}}\right)$on a wide range of formats. The proposed architecture supports both single-path and dual-path execution. The single path is designed for performing acrshort DP acrshort FMA or DPDAC of lower formats, while dual-path supports parallel operations for single-precision (SP) addition and 2-term SP or acrshort TF32 dot-product or 4-term acrshort HP or BF16 dot-product. Moreover, numerical precision conversion is also supported by the proposed architecture, allowing for the conversion of numbers to higher or lower formats. The proposed DPDAC has been demonstrated to significantly reduce the overhead in comparison to discrete designs that utilize multiple single-mode acrshort FP units to achieve the same functionalities. Furthermore, when compared to the state-of-the-art multiple-precision designs, the proposed architecture has been shown to support a wide range of formats and a greater variety of operations with lower costs. Hongbing Tan, Libo Huang 0002, Hui Guo 0004, Qianming Yang, Li Shen 0007, Gang Chen 0023, Liquan Xiao, Nong Xiao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Low-Cost Multiple-Precision Multiplication Unit Design For Deep LearningabstractLow-precision formats have been proposed and applied to deep learning algorithms to speed up training and inference. This paper proposes a novel multiple-precision multiplication unit(MU) for deep learning. The proposed MU supports four types of precision for floating-point(FP) numbers-FP8-E4M3, FP8-E5M2, FP16, FP32-and 8-bit fixed-point(FIX) numbers. The MU can execute four parallel FP8 and eight parallel FIX8 multiplications simultaneously in one cycle, or four parallel FP16 multiplications fully pipelined with a latency of one, or one FP32 multiplication with a latency of one cycle. The simultaneous execution of FIX8 and FP8 can meet the requirements of the specific deep learning algorithms. Thanks to the low-precision-combination(LPC) and vectorization design method, multiplication in any precision can get 100% utilization of the multiplier resources, and the MU can adopt a lower clock delay to achieve better performance in all data types. Compared with the existing multiple-precision units designed for deep learning, this MU can support more types of low-precision formats by lower area overhead; and exhibits higher throughput at FIX8 with at least 8× improvement. Libo Huang 0002, Hongbing Tan, Ling Yang 0008, Qianming Yang |
ACM Great Lakes Symposium on VLSI | 6 |
| 2022 | Efficient Multiple-Precision and Mixed-Precision Floating-Point Fused Multiply-Accumulate Unit for HPC and AI Applications
Hongbing Tan, Run Yan, Ling Yang 0008, Libo Huang 0002, Liquan Xiao, Qianming Yang |
ICA3PP | 6 |
| 2017 | Optimizing OpenCL Implementation of Deep Convolutional Neural Network on FPGA
Yuran Qiao, Junzhong Shen, Dafei Huang, Qianming Yang, Mei Wen, Chunyuan Zhang |
NPC | 4 |
| 2017 | FPGA-accelerated deep convolutional neural networks for high throughput and energy efficiencyabstractSummary Recent breakthroughs in the deep convolutional neural networks (CNNs) have led to great improvements in the accuracy of both vision and auditory systems. Characterized by their deep structures and large numbers of parameters, deep CNNs challenge the computational performance of today. Hardware specialization in the form of field‐programmable gate array offers a promising path towards major leaps in computational performance while achieving high‐energy efficiency. In this paper, we focus on accelerating deep CNNs using the Xilinx Zynq‐zq7045 FPGA SoC. As most of the computational workload can be converted to matrix multiplications, we adopt a matrix multiplier‐based accelerator architecture. Dedicated units are designed to eliminate the conversion overhead. We also design a customized memory system according to the memory access pattern of CNNs. To make the accelerator easily usable by application developers, our accelerator supports Caffe, which is a widely used software framework of deep CNN. Different CNN models can be adopted by our accelerator, with good performance portability. The experimental results show that for a typical application of CNN, image classification, an average throughout of 77.8 GFLOPS is achieved, while the energy efficiency is 4.7× better than an Nvidia K20 GPGPU. © 2016 The Authors. Concurrency and Computation: Practice and Experience Published by John Wiley & Sons Ltd Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen, Chunyuan Zhang |
Concurr. Comput. Pract. Exp. | 4 |
| 2015 | Unified Virtual Memory Support for Deep CNN Accelerator on SoC FPGA
Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen |
ICA3PP (1) | 4 |
| 2013 | Accelerating thread-intensive and explicit memory management programs with dynamic partial reconfiguration
Qianming Yang, Mei Wen, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 1 |
| 2012 | The masala machine: accelerating thread-intensive and explicit memory management programs with dynamically reconfigurable FPGAs (abstract only)abstractA uniform FPGA-based architecture, an efficient programming model and a simple mapping method are paramount for PPGA technology to be more widely accepted. This paper presents MASALA, a dynamically reconfigurable FPGA-based accelerator specifically for parallel programs written in thread-intensive and explicit memory management (TEMM) programming models. The system uses TEMM programming model to parallelize the demanding application, including decomposing the application into separate thread blocks, decoupling compute and data load/store etc. Hardware engines are included into the MASALA by using partial dynamic reconfigure modules, each of which encapsulates Thread Process Engine implementing the thread functionality in hardware. A data dispatching scheme is also included in MASALA to enable the explicit communication among multiple memory hierarchies such as between inter-hardware engines, the host processor and hardware engines. At last, the paper illustrates a Multi-FPGA prototype system of the presented architecture: MASALA-SX. A large synthetic aperture radar (SAR) image formatting experiment shows that the MASALA architecture facilitates the construction of a TEMM program accelerator by providing it with greater performance and less power consumption than current CPU platforms, but without sacrificing programmability, flexibility and scalability. Mei Wen, Nan Wu 0003, Qianming Yang, Chunyuan Zhang |
FPGA | 3 |
| 2010 | Software Managed Instruction Scratchpad Memory Optimization in Stream Architecture Based on Hot Code Analysis of KernelsabstractStream processors, such as Imagine, GPGPUs, FT64 and MASA, typically uses software managed scratchpad instruction memory which improves performance and significantly reduces energy consumption. In this paper, we build a kernel-storage model to analyze the hot spot of kernels in stream programs. Based on the analysis, we define Kernel Hot Code and prove that scratchpad instruction memory should focus on the access efficiency of it. A methodology for finding Kernel Hot Code in the kernels of different structures is presented as well. In accordance with this method, we develop HOIS for Stream Architecture, which adopts a software managed scratchpad memory to store Kernel Hot Code, and uses a small hardware managed victim cache to store the Kernel Cool Code. HOIS is evaluated by measuring the performance of six applications on the MASA_S simulation platform. The results show that HOIS can achieve high efficiency in predictable applications with little performance loss. Yi He 0008, Ju Ren 0002, Mei Wen, Qianming Yang, Nan Wu 0003, Chunyuan Zhang |
DSD | 4 |
| 2007 | FT64: Scientific Computing with Streams
Mei Wen, Nan Wu 0003, Chunyuan Zhang, Qianming Yang, Changqing Xun |
HiPC | 5 |