VLDB 2026 Research / reviewers in the wild / expert
Hongbing Tan
dblp:203/9321
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | X-SA: An Efficient Configurable Systolic Array Computing Architecture for GPGPUabstractGPGPUs are pivotal for edge AI, but resource constraints demand efficient low-precision computation. Conventional GPGPUs face challenges in resource utilization, particularly with irregular matrices common in AI, and memory bandwidth limitations on edge devices. Traditional fixed-size systolic arrays often suffer from underutilization under varying workloads. This paper introduces X-SA, a configurable systolic array architecture tailored for INT8 matrix multiplication on GPGPUs in resource-constrained edge environments. X-SA distinctively employs a parameterized$2 \times N$processing element design enabling dynamic computational scaling, unlike fixed systolic arrays. It integrates an interleaved matrix buffer to alleviate memory bottlenecks and optimize dataflow. Experimental results demonstrate X-SA achieves a$2.83 \times$performance speedup over the Vortex baseline with minimal Look-Up Table overhead of 2.8% and Flip-Flops overhead of 1.4%.. It offers comparable performance to a standard$4 \times 4$systolic array but with significantly reduced area by 46.26% and power by 39.42%, and superior processing element utilization for irregular matrices. X-SA provides a approach to help improve the performance of some AI applications running on edge GPGPUs relatively in resource-constrained environments. Yingsong Wang, Zhenzhen Jia, Ling Yang 0008, Hongbing Tan, Junsheng Chang, Junbo Tie, Libo Huang 0002 |
HPCC | 4 |
| 2025 | PolyPE: An Efficient Multi-Precision Multi-Mode Floating-Point Processing Element for HPC and AIabstractIn this paper, an efficient multi-precision multimode floating-point Processing Element is designed for HPCenabled AI workloads, called PolyPE, in which Poly means multiprecision multi-mode. It supports both conventional and mixedprecision FMA operations, including single-FMA, dual-FMA, and quad-FMA modes, as well as quad-FMA-add for enhanced throughput. The supported precisions include double precision, single precision, half precision, TF32, and BF16. At each clock cycle, the processing element can perform one double-precision, two single-precision, or four half-precision operations. Compared to existing designs, it offers broader precision support, including TF32 and BF16, with higher throughput and lower hardware overhead, achieving up to 5× improvement over standard FMA. We integrated the design into an open-source GPGPU and extended its instruction set. Experimental results show up to 2.17× performance gain, with 27.2% and 41.2% reductions in LUT and FF usage, respectively, while preserving functional equivalence. Zhenzhen Jia, Hongbing Tan, Ling Yang 0008, Hui Guo 0004, Junsheng Chang, Yongwen Wang, Libo Huang 0002 |
ICCD | 2 |
| 2024 | ImSPU: Implicit Sharing of Computation Resources Between Vector and Scalar Processing Units
Hongbing Tan, Guichu Sun, Liquan Xiao, Yuanhu Cheng, Quan Deng 0003, Bingcai Sui, Yongwen Wang, Libo Huang 0002 |
Euro-Par (2) | 1 |
| 2024 | A Low-Cost Floating-Point Dot-Product-Dual-Accumulate Architecture for HPC-Enabled AIabstractThe dot-product$\sum _{i=1}^{N} A_{i}\times B_{i}$is one of the most frequently used operations for a wide variety of high-performance computing (HPC) and artificial intelligence (AI) applications. However, for large-scale algorithms, such as acrshort GEMM and acrshort FFT, independent additions are necessary to accumulate the results of length-limited dot-product in order to form the final result, thus increasing latency and overhead. Hence, we proposed a dot-product-dual-accumulate (DPDAC) architecture capable of performing$\left({\sum _{i=1}^{N=1,2,4} A_{i}\times B_{i} + \sum _{j=1}^{M=1,2} C_{j}}\right)$on a wide range of formats. The proposed architecture supports both single-path and dual-path execution. The single path is designed for performing acrshort DP acrshort FMA or DPDAC of lower formats, while dual-path supports parallel operations for single-precision (SP) addition and 2-term SP or acrshort TF32 dot-product or 4-term acrshort HP or BF16 dot-product. Moreover, numerical precision conversion is also supported by the proposed architecture, allowing for the conversion of numbers to higher or lower formats. The proposed DPDAC has been demonstrated to significantly reduce the overhead in comparison to discrete designs that utilize multiple single-mode acrshort FP units to achieve the same functionalities. Furthermore, when compared to the state-of-the-art multiple-precision designs, the proposed architecture has been shown to support a wide range of formats and a greater variety of operations with lower costs. Hongbing Tan, Libo Huang 0002, Hui Guo 0004, Qianming Yang, Li Shen 0007, Gang Chen 0023, Liquan Xiao, Nong Xiao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | A Multi-level Parallel Integer/Floating-Point Arithmetic Architecture for Deep Learning Instructions
Hongbing Tan, Libo Huang 0002, Dezun Dong, Yongwen Wang, Liquan Xiao |
Euro-Par | 1 |
| 2023 | Low-Cost Multiple-Precision Multiplication Unit Design For Deep LearningabstractLow-precision formats have been proposed and applied to deep learning algorithms to speed up training and inference. This paper proposes a novel multiple-precision multiplication unit(MU) for deep learning. The proposed MU supports four types of precision for floating-point(FP) numbers-FP8-E4M3, FP8-E5M2, FP16, FP32-and 8-bit fixed-point(FIX) numbers. The MU can execute four parallel FP8 and eight parallel FIX8 multiplications simultaneously in one cycle, or four parallel FP16 multiplications fully pipelined with a latency of one, or one FP32 multiplication with a latency of one cycle. The simultaneous execution of FIX8 and FP8 can meet the requirements of the specific deep learning algorithms. Thanks to the low-precision-combination(LPC) and vectorization design method, multiplication in any precision can get 100% utilization of the multiplier resources, and the MU can adopt a lower clock delay to achieve better performance in all data types. Compared with the existing multiple-precision units designed for deep learning, this MU can support more types of low-precision formats by lower area overhead; and exhibits higher throughput at FIX8 with at least 8× improvement. Libo Huang 0002, Hongbing Tan, Ling Yang 0008, Qianming Yang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | A Scalable BFloat16 Dot-Product Architecture for Deep LearningabstractBFloat16(BF16) format has recently driven the development of deep learning due to its higher energy efficiency and less memory consumption than the traditional format. This paper presents a scalable BF16 dot-product(DoP) architecture for high-performance deep-learning computing. A novel 4-term DoP unit is proposed as a fundamental module in the architecture, which performs 4-term DoP operation in three cycles. More-term DoP units are constructed through the extension of the fundamental unit, in which early exponent comparison is performed to hide latency, and intermediate normalization and rounding are omitted to improve accuracy and further reduce latency. Compared with the discrete design, the proposed architecture reduces latency by 22.8% for 4-term DoP, and a larger proportion of latency is reduced as the size of the DoP operation increases. Compared with existing designs for BF16, the proposed architecture at 64-term exhibits better-normalized energy efficiency and higher throughput with at least 1.88× and 20.3× improvement, respectively. Libo Huang 0002, Hongbing Tan, Hui Guo 0004 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | SFDoP: A Scalable Fused BFloat16 Dot-Product Architecture for DNNabstractThe BFloat16(BF16) format has emerged as a driving force in Deep Neural Networks(DNNs), owing to its superior energy efficiency and lower memory footprint than traditional formats. Since the BF16 format is mainly used in computation-intensive layers such as the general matrix multi-plication(GEMM) layer, this paper presents SFDoP, a scalable BF16 fused dot-product(DoP) architecture for high-performance computation in DNNs. The SFDoP features a novel fused 4-term DoP unit as a basic unit, which performs 4-term DoP operation in three cycles. More-term DoP units are constructed by extending this basic unit. The extended units incorporate early exponent comparison to mask latency and omit intermediate normalization and rounding to further improve performance. Compared with discrete designs, SFDoP-4 reduces latency by 15.6% for 4-term DoP operation, with greater reductions achieved in the extended units. Compared with existing BF16 designs, SFDoP exhibits improved throughput and energy efficiency, with gains of at least 82.2% and 28.1%, respectively. For GEMM operation of large size, SFDoP achieves better performance in the extended units than the basic unit. Hongbing Tan, Libo Huang 0002 |
ICCD | 2 |
| 2023 | Multiple-Mode-Supporting Floating-Point FMA Unit for Deep Learning ProcessorsabstractIn this article, a new multiple-mode floating-point fused multiply–add (FMA) unit is proposed for deep learning processors. The proposed design supports three functional modes—normal FMA mode, mixed FMA mode, and dual FMA mode—and four types of precision—single-precision (SP), half-precision (HP), BFloat16 (BF16), and TensorFloat-32 (TF32)—based on the practical requirements of deep learning applications. In the normal FMA mode, conventional FMA operations, one SP operation or two parallel HP operations, are performed every clock cycle. In the mixed FMA mode and dual FMA mode, mixed-precision operations, the fused multiply–accumulate and the dot-product, are implemented, respectively. Specifically, the product of lower precision multiplication can be accumulated to a higher precision addend. Compared with the mixed FMA mode, the throughput is doubled in the dual FMA mode due to the full utilization of the multiplier operand bandwidth. In addition to FMA operations, numerical precision conversion (NPCvt) is also supported in this work: higher precision FMA results can be converted into lower precision numbers, corresponding to the datatype transform in the datapath of deep neural network (DNN) training. The FMA design presented herein uses both the segmentation and reusing methods to trade off performance, such as throughput and latency, against area, and power. Compared with the state-of-the-art multiple-precision FMA unit, the proposed design supports more types of floating-point operation and NPCvt, with higher throughput and lower hardware overhead. Hongbing Tan, Gan Tong, Libo Huang 0002, Liquan Xiao, Nong Xiao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Efficient Multiple-Precision and Mixed-Precision Floating-Point Fused Multiply-Accumulate Unit for HPC and AI Applications
Hongbing Tan, Run Yan, Ling Yang 0008, Libo Huang 0002, Liquan Xiao, Qianming Yang |
ICA3PP | 1 |
| 2017 | Modeling and evaluation for gather/scatter operations in Vector-SIMD architecturesabstractGather/scatter are state of the art vector memory access modes in Vector-SIMD architectures. However, because of the stochastic and complicated properties, the hardware design of gather/scatter operations lacks theoretical analysis and modeling. This paper proposes a model for gather/scatter operations on local vector memory for the first time. The model can not only give all the possible distributions of access locations, calculate the probability of access conflicts and predict the number of access conflicts, but also can provide the theoretical guidance for the performance optimization. This model is validated through experiments which can guide users to more specifically design and optimize the implementation of gather/scatter operations. Hongbing Tan, Sheng Liu 0001 |
ASAP | 1 |