Hao Zhang 0041

dblp:55/2270-41 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-3027-0485ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 When Posit Meets Microscaling: Energy Efficient Posit-Based Processing Element for Edge AI Computation
abstract
Low-precision computation is an effective method to improve energy efficiency when processing AI models at the edge. The design of numeric format is important to maintain good accuracy while reducing energy consumption. However, current fixed-point based formats or floating-point based formats have either limitations in representation range or precision, and thus efficiency or accuracy is compromised. Posit formats can achieve both large dynamic range and high precision, however, the computation overhead is too high. Inspired by the recent microscaling format, in this paper, a novel microscaling posit format and its corresponding dot-product based processing element are proposed. By designing a specific format, the dotproduct computation overhead of the original posit format is significantly reduced. Implementation results show that the proposed processing element can achieve up to $79 \%$ area reduction and $74 \%$ power reduction when compared with other designs available in the literature, which makes the proposed designs especially suitable for edge AI computation.
Seok-Bum Ko, Zhiqiang Wei 0002, Hao Zhang 0041
ASP-DAC5
2026 Power-Efficient and Reconfigurable Compute Unit for Multi-Precision AI Inference at the Edge
Muhammad Hamis Haider, Hao Zhang 0041, Seok-Bum Ko
ISCAS2
2026 FlexPWL: A Flexible, Scalable, and Multiplier-Free Approach for Activation Functions on FPGA
abstract
The hardware implementation of nonlinear activation functions (AFs), such as the Sigmoid and Hyperbolic Tangent (Tanh), presents a significant bottleneck for deploying Recurrent Neural Networks (RNNs) on resource-constrained edge devices. Field Programmable Gate Arrays (FPGAs), as a leading platform for edge AI, face challenges in efficiently executing these functions due to their complex mathematical nature and limited computational resources. This paper proposes a novel, hardware-friendly piecewise linear (PWL) approximation method for implementing Sigmoid and Tanh functions. Our approach introduces new mathematical formulations for computing the y-intercept that significantly reduce Mean Squared Error (MSE) and Maximum Absolute Error (MXE) at no additional hardware cost. By consolidating the features of recent works, the proposed architecture eliminates a costly segment address encoder, supports pipelining for low-latency and high-frequency operation, and avoids the complexity and low precision of other methods. Moreover, it offers high configurability in bit width, number of segments, and input ranges, enabling adaptable deployment across diverse hardware and precision targets. A unified architecture is presented for both AFs, maintaining identical FPGA resource usage across functions. Experimental results demonstrate that the proposed method outperforms prior state-of-the-art designs by up to 18.52× in accuracy, 437.29× in latency, and 3.71× in frequency, and achieves LUT and flip-flop savings of up to 11.63× and 12.42×, respectively.
Ebrahim Fard, Janier Arias-Garcia, Hao Zhang 0041, Seok-Bum Ko
IEEE Trans. Computers3
2026 T3: Transformer Accelerator With Efficient Top-K Sorting and Dynamic Tanh Computation for Marine Edge Computing
abstract
The Transformer model has demonstrated superior performance across numerous marine applications. However, the complexity of self-attention and layer normalization in addition to the high precision requirement poses challenges for their efficient deployment in marine edge devices. To address this issue, in this article, an FPGA-based Transformer accelerator, T3, is proposed for marine edge computing. Two main techniques are proposed and utilized in the proposed T3accelerator. The first technique is the design of a floating-point (FP)-based top-$K$sorting method for self-attention pruning, which can be used to reduce the computational cost of self-attention modules. The other technique is an efficient implementation of the dynamic Tanh (DyT) module, which utilizes error-controlled piecewise linear (PWL) approximation and coefficient fusion method, which can be used to take the place of the costly layer normalization. The two proposed modules and the whole T3accelerator are implemented in Xilinx UltraScale+ FPGA devices. The proposed top-$K$sorting method can achieve up to 79.6% reduction in lookup tables (LUTs), 80.3% reduction in flip-flops (FFs), and 90.8% reduction in power consumption when compared with the previous top-$K$sorting method. The proposed DyT module can consume 38.3% fewer LUTs while achieving better accuracy compared to the state-of-the-art designs. Finally, due to the effectiveness of the proposed techniques, the proposed T3accelerator can achieve 15.2% higher energy efficiency while consuming fewer number of logic resources.
Dingyang Yu, Changlong Chen, Seok-Bum Ko, Hao Zhang 0041
IEEE Trans. Very Large Scale Integr. Syst.5
2024 Energy Efficient FPGA-Based Binary Transformer Accelerator for Edge Devices
abstract
Transformer-based large language models have gained much attention recently. Due to their superior performance, they are expected to take the place of conventional deep learning methods in many fields of applications, including edge computing. However, transformer models have even more amount of computations and parameters than convolutional neural networks which makes them challenging to be deployed at resource-constrained edge devices. To tackle this problem, in this paper, an efficient FPGA-based binary transformer accelerator is proposed. Within the proposed architecture, an energy efficient matrix multiplication decomposition method is proposed to reduce the amount of computation. Moreover, an efficient binarized Softmax computation method is also proposed to reduce the memory footprint during Softmax computation. The proposed architecture is implemented on Xilinx Zynq Untrascale+ device and implementation results show that the proposed matrix multiplication decomposition method can reduce up to 78% of computation at runtime. The proposed transformer accelerator can achieve improved throughput and energy efficiency compared to previous transformer accelerator designs.
Congpeng Du, Seok-Bum Ko, Hao Zhang 0041
ISCAS3
2024 Anterior mediastinal nodular lesion segmentation from chest computed tomography imaging using UNet based neural network with attention mechanisms
Yi Wang 0064, Won Gi Jeong, Hao Zhang 0041, Younhee Choi, Gong Yong Jin, Seok-Bum Ko
Multim. Tools Appl.3
2024 Decoder Reduction Approximation Scheme for Booth Multipliers
abstract
Existing approximate Booth multipliers fail to keep up with modern approximate multipliers such as truncation-based approximate logarithmic multipliers. This paper introduces a new approximation scheme for Booth multipliers that can operate with negligible error rates using only$N/4$Booth decoders, instead of the traditional$N/2$Booth decoders. The proposed 16-bit BD16.4 approximate Booth multiplier reduces the Normalized Mean Error Deviation (NMED) by 96.5% and the Power-Area-Product (PAP) by 69.6%, when compared to a state-of-the-art approximate logarithmic multiplier. Additionally, the proposed BD16.4 approximate multiplier reduces the NMED by 94.4% and PAP by 74.8%, when compared to a state-of-the-art higher-radix approximate Booth multiplier. The proposed 8-bit approximate Booth multipliers reduce the NMED by up to 74% and PAP by up to 5% when compared to the existing state-of-the-art approximate logarithmic multipliers. We validated the results derived in this paper through a neural network inference experiment, where the proposed approximate multipliers showed a negligible drop in inference accuracy compared to the exact Booth multipliers and the state-of-the-art approximate logarithmic multipliers (ALM). The proposed approximate multipliers achieved a Power-Delay-Product reduction of 63% (vs. exact) and 21.22% (vs. ALM) in 16-bit experiments and a reduction of 67% (vs. exact) and 8.75% (vs. ALM) in 8-bit experiments.
Muhammad Hamis Haider, Hao Zhang 0041, Seok-Bum Ko
IEEE Trans. Computers2
2024 Merge Loss Calculation Method for Highly Imbalanced Data Multiclass Classification
abstract
In real classification scenarios, the number distribution of modeling samples is usually out of proportion. Most of the existing classification methods still face challenges in comprehensive model performance for imbalanced data. In this article, a novel theoretical framework is proposed that establishes a proportion coefficient independent of the number distribution of modeling samples and a general merge loss calculation method independent of class distribution. The loss calculation method of the imbalanced problem focuses on both the global and batch sample levels. Specifically, the loss function calculation introduces the true-positive rate (TPR) and the false-positive rate (FPR) to ensure the independence and balance of loss calculation for each class. Based on this, global and local loss weight coefficients are generated from the entire dataset and batch dataset for the multiclass classification problem, and a merge weight loss function is calculated after unifying the weight coefficient scale. Furthermore, the designed loss function is applied to different neural network models and datasets. The method shows better performance on imbalanced datasets than state-of-the-art methods.
Zehua Du, Hao Zhang 0041, Zhiqiang Wei 0002, Xianqing Huang
IEEE Trans. Neural Networks Learn. Syst.2
2022 Segmentation for document layout analysis: not dead yet
Logan Markewich, Hao Zhang 0041, Yubin Xing, Navid Lambert-Shirzad, Zhexin Jiang, Roy Ka-Wei Lee, Seok-Bum Ko
Int. J. Document Anal. Recognit.2
2021 Efficient Multiple-Precision Posit Multiplier
abstract
Posit number system has been recently widely applied in many fields of applications. For different applications, the precision requirements are usually different. In addition, the transprecision computing paradigm, which is proposed for energy efficient computation, even requires different precision in each computation step. To support computations of various precision in a single hardware architecture, in this paper, a unified architecture of multiple-precision posit multiplier is proposed. The proposed posit multiplier supports the commonly used Posit(8, 0), Posit(16, 1), and Posit(32, 2) formats, where one Posit(32, 2), or two parallel Posit(16, 1), or four parallel Posit(8, 0) multiplications can be accomplished each time. Each module of the proposed posit multiplier is carefully tailored for resource sharing among three supported precision formats. Compared to the Posit(32, 2) multiplier, the proposed multiple- precision multiplier adds the support for parallel low-precision posit multiplications with only 12.8% more area and 15.4% more power. The proposed architecture can be used in posit-enabled general-purpose processor designs.
Hao Zhang 0041, Seok-Bum Ko
ISCAS1
2021 Energy efficient spiking neural network processing using approximate arithmetic units and variable precision weights
Yi Wang 0064, Hao Zhang 0041, Kwang-Il Oh, Jae-Jin Lee, Seok-Bum Ko
J. Parallel Distributed Comput.2
2021 A Real-Time Architecture for Pruning the Effectual Computations in Deep Neural Networks
abstract
Integrating Deep Neural Networks (DNNs) into the Internet of Thing (IoT) devices could result in the emergence of complex sensing and recognition tasks that support a new era of human interactions with surrounding environments. However, DNNs are power-hungry, performing billions of computations in terms of one inference. Spatial DNN accelerators in principle can support computation-pruning techniques compared to other common architectures such as systolic arrays. Energy-efficient DNN accelerators skip bit-wise or word-wise sparsity in the input feature maps (ifmaps) and filter weights which means ineffectual computations are skipped. However, there is still room for pruning the effectual computations without reducing the accuracy of DNNs. In this paper, we propose a novel real-time architecture and dataflow by decomposing multiplications down to the bit level and pruning identical computations in spatial designs while running benchmark networks. The proposed architecture prunes identical computations by identifying identical bit values available in both ifmaps and filter weights without changing the accuracy of benchmark networks. When compared to the reference design, our proposed design achieves an average per layer speedup of$\times 1.4$and an energy efficiency of$\times 1.21$per inference while maintaining the accuracy of benchmark networks.
Mohammadreza Asadikouhanjani, Hao Zhang 0041, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference
abstract
In this paper, a new flexible multiple-precision multiply-accumulate (MAC) unit is proposed for deep neural network training and inference. The proposed MAC unit supports both fixed-point operations and floating-point operations. For floating-point format, the proposed unit supports one 16-bit MAC operation or sum of two 8-bit multiplications plus a 16-bit addend. To make the proposed MAC unit more versatile, the bit-width of exponent and mantissa can be flexibly exchanged. By setting the bit-width of exponent to zero, the proposed MAC unit also supports fixed-point operations. For fixed-point format, the proposed unit supports one 16-bit MAC or sum of two 8-bit multiplications plus a 16-bit addend. Moreover, the proposed unit can be further divided to support sum of four 4-bit multiplications plus a 16-bit addend. At the lowest precision, the proposed MAC unit supports accumulating of eight 1-bit logic AND operations to enable the support of binary neural networks. Compared to the standard 16-bit half-precision MAC unit, the proposed MAC unit provides more flexibility with only 21.8 percent area overhead. Compared to a standard 32-bit single-precision MAC unit, the proposed MAC unit requires much less hardware cost but still provides 8-bit exponent in the numerical format to maintain large dynamic range for deep learning computing.
Hao Zhang 0041, Dongdong Chen 0002, Seok-Bum Ko
IEEE Trans. Computers1
2019 Efficient Posit Multiply-Accumulate Unit Generator for Deep Learning Applications
abstract
The recently proposed posit number system is more accurate and can provide a wider dynamic range than the conventional IEEE754-2008 floating-point numbers. Its nonuniform data representation makes it suitable in deep learning applications. Posit adder and posit multiplier have been well developed recently in the literature. However, the use of posit in fused arithmetic unit has not been investigated yet. In order to facilitate the use of posit number format in deep learning applications, in this paper, an efficient architecture of posit multiply-accumulate (MAC) unit is proposed. Unlike IEEE754-2008 where four standard binary number formats are presented, the posit format is more flexible where the total bitwidth and exponent bitwidth can be any number. Therefore, in this proposed design, bitwidths of all datapath are parameterized and a posit MAC unit generator written in C language is proposed. The proposed generator can generate Verilog HDL code of posit MAC unit for any given total bitwidth and exponent bitwidth. The code generated by the generator is a combinational design, however a 5-stage pipeline strategy is also presented and analyzed in this paper. The worst case delay, area, and power consumption of the generated MAC unit under STM-28nm library with different bitwidth choices are provided and analyzed.
Hao Zhang 0041, Jiongrui He, Seok-Bum Ko
ISCAS1
2019 Efficient Multiple-Precision Floating-Point Fused Multiply-Add with Mixed-Precision Support
abstract
In this paper, an efficient multiple-precision floating-point fused multiply-add (FMA) unit is proposed. The proposed FMA supports not only single-precision, double-precision, and quadruple-precision operations, as some previous works do, but also half-precision operations. The proposed FMA architecture can execute one quadruple-precision operation, or two parallel double-precision operations, or four parallel single-precision operations, or eight parallel half-precision operations every clock cycle. In addition to the support of normal FMA operations, the proposed FMA also supports mixed-precision FMA operations and mixed-precision dot-product operations. Specifically, the products of two lower precision multiplications can be accumulated to a higher precision addend. By setting the operands of one multiplication to zeros, the proposed FMA can also perform mixed-precision FMA operations. Support for mixed-precision FMA and mixed-precision dot-product is newly added but it only consumes 6.5 percent more area compared to a normal multiple-precision FMA unit. Compared to the state-of-the-art multiple-precision FMA design, the proposed FMA supports more floating-point operations such as half-precision FMA operations and mixed-precision operations with only 10.6 percent larger area.
Hao Zhang 0041, Dongdong Chen 0002, Seok-Bum Ko
IEEE Trans. Computers1
2018 Efficient Fixed/Floating-Point Merged Mixed-Precision Multiply-Accumulate Unit for Deep Learning Processors
abstract
Deep learning is getting more and more attentions in recent years. Many hardware architectures have been proposed for efficient implementation of deep neural network. The arithmetic unit, as a core processing part of the hardware architecture, can determine the functionality of the whole architecture. In this paper, an efficient fixed/floating-point merged multiply-accumulate unit for deep learning processor is proposed. The proposed architecture supports 16-bit half-precision floating-point multiplication with 32-bit single-precision accumulation for training operations of deep learning algorithm. In addition, within the same hardware, the proposed architecture also supports two parallel 8-bit fixed-point multiplications and accumulating the products to 32-bit fixed-point number. This will enable higher throughput for inference operations of deep learning algorithms. Compared to a half-precision multiply-accumulate unit (accumulating to single-precision), the proposed architecture has only 4.6% area overhead. With the proposed multiply-accumulate unit, the deep learning processor can support both training and high-throughput inference.
Hao Zhang 0041, Seok-Bum Ko
ISCAS1