EDBT 2026 Demo / reviewers in the wild / expert
Hongbing Pan
dblp:64/9429
· DBLP profile ↗
19ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-7181-8278ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TEA-M: A Tiny and Efficient Architecture for Multivalue Image Watermarking With Versatile Hardware ImplementationabstractThe evolution of modern communication and the Internet has enhanced the accessibility, editability, and dissemination speed of the digital images, intensifying the need for digital image copyright protection. Digital watermarking technique is widely applied to the copyright protection of digital images. Compared with binary watermarking, multi-value watermarking can convey richer copyright information while offering higher robustness. However, existing multi-value watermarking algorithms suffer from high computational complexity, which makes it difficult to meet the ever-increasing demands for real-time and high-speed images or videos processing tasks. In this paper, we propose the TEA-M, a tiny and efficient multi-value watermarking architecture. By applying a multiplier-free approximate discrete cosine transform (DCT) transform to the Z channel and leveraging two novel multi-value watermarking strategies based on remainder adjustment, TEA-M achieves up to 49× higher energy efficiency and 310× higher area efficiency while maintaining high throughput compared with the latest approaches. For the FPGA implementation, TEA-M can achieve 5007 frames/s at its highest operating frequency of 328.299 MHz, with a peak energy efficiency of 4.146 × 105Mbps/W. For the ASIC implementation, TEA-M achieves 7629 frames/s at 500 MHz, with a peak energy efficiency of 5.36×106Mbps/W and an area efficiency of 2.1×106Mbps/mm2. In practical applications, TEA-M further offers users diverse watermarking configuration options to meet the needs of various scenarios. Zikang Wang, Zhengyu Mei, Feng Yan 0002, Hongbing Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | PIMapping: A Tile-Level Dataflow Optimization Framework for PIM ArchitectureabstractProcess-in-memory (PIM) accelerators demonstrate outstanding performance in accelerating matrix-vector multiplication (MVM) tasks in neural networks. To achieve better acceleration performance, extensive research has focused on the design of hierarchical tile-based PIM architectures, presenting challenges for the hardware deployment of algorithms. Multilayer parallelism in tile-based architectures requires the support of dataflow optimization techniques. However, existing research primarily performs dataflow analysis using computational models and performance metrics that are not suitable for tile-level dataflows. In this paper, we propose an analytical framework, PIMapping, for tile-based PIM architectures that supports multi-layer parallelism mapping and tile-level dataflow optimization. In this work, we first establish a general dataflow representation for tile-level dataflow, which serves as an intermediate representation (IR) for multi-layer DNN mapping and scheduling tasks at tile level. Next, based on the proposed representation, we introduce a data proximity-based mapping method aimed at minimizing the inter-tile communication overhead. Furthermore, we propose a congestion-aware scheduling algorithm to minimize inter-tile communication conflicts. Experimental case-studies are conducted to map common DNN algorithms onto different PIM architectures. The results demonstrate significant improvements over the state-of-the-art mapping framework in terms of communication latency, required inter-tile bandwidth, and pipeline efficiency. Ziqian Zhu, Jinsen Zhu, Hongbing Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | TEA-SPS: A Tiny and Efficient Architecture for Softmax With Parallelism and Sparsity AdaptabilityabstractWith the remarkable performance of Transformer-based networks in multiple fields and increasing demand for computational resources by softmax within them, it is inevitable for hardware accelerators to support softmax in Transformer. However, due to the lack of co-design of algorithm and hardware, there still remains space for optimizing the hardware architecture for softmax. Therefore, TEA-SPS is proposed as an algorithm and hardware co-designed architecture to improve softmax with two methods: Configurable Parallelism softmax with Sparse mask Strategy (CPSS) and Specific Piecewise Information Extractor (SPIE). CPSS has the advantage of supporting different throughput requirements through configurable data-level parallelism and performing sparse masking on the outputs to reduce the computational load and memory access of subsequent operations. To further explore the optimal solution set among the design space of parameters in CPSS, SPIE is proposed to achieve co-optimization of accuracy and hardware overhead. Based on them, the efficient hardware architecture of TEA-SPS is proposed. The implementation results show that at the frequency of 0.5 GHz under TSMC 90-nm technology, the peak efficiency of TEA-SPS processing 8-bit quantized data can reach up to 216.97 Gps/(mm$\boldsymbol {^{2}\cdot }$mW), with the area of$3290.21~\boldsymbol {\mu }$m$\boldsymbol {^{2}}$and the power consumption of 0.7004 mW. In addition, TEA-SPS provides support for input sequences of arbitrary length with negligible accuracy loss compared to the quantized baseline, while achieving an average sparse rate of 68.6% on the GLUE tasks. Zhanhao Cui, Zhengyu Mei, Feng Yan 0002, Hongbing Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | RDIF: Infrared and Visible Image Fusion Based on Reverse Cross-Attention and Diffusion Model
Hongli Su, Yuchen Hong, Chenglei Peng, Hongbing Pan |
ICANN (2) | 5 |
| 2025 | ONNXPruner: ONNX-Based General Model Pruning AdapterabstractRecent advancements in model pruning have focused on developing new algorithms and improving upon benchmarks. However, the practical application of these algorithms across various models and platforms remains a significant challenge. To address this challenge, we propose ONNXPruner, a versatile pruning adapter designed for the ONNX format models. ONNXPruner streamlines the adaptation process across diverse deep learning frameworks and hardware platforms. A novel aspect of ONNXPruner is its use of node association trees, which automatically adapt to various model architectures. These trees clarify the structural relationships between nodes, guiding the pruning process, particularly highlighting the impact on interconnected nodes. Furthermore, we introduce a tree-level evaluation method. By leveraging node association trees, this method allows for a comprehensive analysis beyond traditional single-node evaluations, enhancing pruning performance without the need for extra operations. Experiments across multiple models and datasets confirm ONNXPruner's strong adaptability and increased efficacy. Our work aims to advance the practical application of model pruning. Dongdong Ren, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Jing Huo, Hongbing Pan, Yang Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | MBS: A High-Precision Approximation Method for Softmax and Efficient Hardware ImplementationabstractThe softmax function needs to be frequently used in the multi-head attention layer of Transformer networks. Compared to DNNs and other networks, Transformers have higher computational complexity, requiring higher accuracy and hardware performance for softmax function calculations. Therefore, we propose mixed-base softmax (MBS) for the first time for the approximation of the softmax function. This method combines exponential functions with bases of 2 and 4, which is advantageous for hardware implementation. MBS has a high similarity to the softmax function and demonstrates advanced performance during inference in Transformer network. Through algorithm transformation and hardware optimization, we have designed a low-complexity and highly parallel hardware architecture, which only occupies few additional hardware resources compared to base-2 softmax but achieves higher accuracy. Experimental results show that, under TSMC 90nm CMOS technology at the frequency of 0.5 GHz, our design can achieve the efficiency of 236.18 Gps/(mm2⋅mW) with the area of 4234 μm2. Furthermore, MBS exhibits higher computational accuracy and inference precision compared with base-2 softmax. Yuanchen Wu, Zhiheng Xie, Hongbing Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Leveraging Frequency Analysis for Image Denoising Network PruningabstractAs a common model compression technique, network pruning is widely used to reduce storage and computational cost of deep models in the resource-constrained regime. However, most current pruning methods are designed for high-level vision tasks, with few developed for low-level vision tasks. We observed that the norm-based pruning criterion, originally designed for high-level vision tasks, is highly unsuitable for low-level image denoising networks. This difference arises because image denoising networks pursue distinct feature granularities and goals compared to typical high-level vision tasks. To address this issue, we propose a novel filter evaluation method, termed High-Frequency Components Pruning (HFCP), specifically tailored for image denoising network pruning. HFCP assesses filter importance based on high-frequency components. To the best of our knowledge, this is the first pruning method designed specifically for image denoising tasks, straightforward and applicable to various types of noise. Furthermore, HFCP enhances the pruned model's high-frequency information content with high reliability and interpretability. This facilitates the network's ability to distinguish high-frequency signals from noise. We comprehensively analyzed multiple image denoising networks and validated HFCP's effectiveness across four mainstream networks. Dongdong Ren, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Hongbing Pan, Yang Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | SHINE: protein language model-based pathogenicity prediction for short inframe insertion and deletion variantsabstractAccurate variant pathogenicity predictions are important in genetic studies of human diseases. Inframe insertion and deletion variants (indels) alter protein sequence and length, but not as deleterious as frameshift indels. Inframe indel Interpretation is challenging due to limitations in the available number of known pathogenic variants for training. Existing prediction methods largely use manually encoded features including conservation, protein structure and function, and allele frequency to infer variant pathogenicity. Recent advances in deep learning modeling of protein sequences and structures provide an opportunity to improve the representation of salient features based on large numbers of protein sequences. We developed a new pathogenicity predictor for SHort Inframe iNsertion and dEletion (SHINE). SHINE uses pretrained protein language models to construct a latent representation of an indel and its protein context from protein sequences and multiple protein sequence alignments, and feeds the latent representation into supervised machine learning models for pathogenicity prediction. We curated training data from ClinVar and gnomAD, and created two test datasets from different sources. SHINE achieved better prediction performance than existing methods for both deletion and insertion variants in these two test datasets. Our work suggests that unsupervised protein language models can provide valuable information about proteins, and new methods based on these models can improve variant interpretation in genetic analyses. Hongbing Pan, Alan Tian, Wendy K. Chung, Yufeng Shen |
Briefings Bioinform. | 2 |
| 2022 | TEA-Z: A Tiny and Efficient Architecture Based on Z Channel for Image Watermarking and Its Versatile Hardware ImplementationabstractWith the popularity of smart phones and digital cameras, authentication and protection of digital images have become important to every producer and consumer. Although many digital watermarking methods have been proposed to mitigate the growing problem of online piracy, their efficient deployment for real-time processing in hardware lacks sufficient attention. Hence, with the first introduction of${Z}$channel, this article advances a tiny and efficient architecture named TEA-Z to implement digital watermarking for color images. Avoiding complex computing units with the general color space conversion (GCSC) and simplified discrete cosine transform (DCT), TEA-Z can better achieve the design goals of low complexity, high efficiency, and easy deployment compared with the latest methods. For FPGA implementation, TEA-Z can achieve the maximum frame rate of 3150 frames/s for 256$\times $256 color images and peak energy efficiency of 4.13$\times \,\,\textbf {10}^{\textbf {5}}$Mbps/W at the maximum frequency of 206.48 MHz. For ASIC implementation, the peak energy efficiency will be 2.40$\times \,\,\textbf {10}^{\textbf {7}}$Mbps/W with the tiny area of 0.00356$\textbf {mm}^{\textbf {2}}$and ultralow power of 0.2002 mW at the frequency of 200 MHz under 90-nm CMOS technology. Moreover, TEA-Z can provide various watermarking strategies with its configurations to meet different needs. Zhengyu Mei, Hongbing Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Ultralow-Latency VLSI Architecture Based on a Linear Approximation Method for Computing Nth Roots of Floating-Point NumbersabstractState-of-the-art approaches that perform root computations based on the COordinate Rotation Digital Computer (CORDIC) algorithm suffer from high latency in performing multiple iterations. Therefore, root computations based on the CORDIC algorithm cannot meet the strict latency requirements of some applications. In this paper, we propose a methodology for performing Nth root computations on floating-point numbers based on the piecewise linear (PWL) approximation method. The proposed method divides an Nth root computation into several subtasks approximated by the PWL algorithm. It determines the widest segments of the subtasks and the smallest fractional width needed to satisfy the predefined maximum relative error Max_Errr. Our design is coded in Verilog HDL and synthesized under TSMC 40 nm CMOS technology. The synthesized results show that our design can reach the highest frequency of 2.703 GHz with an area consumption of 2608.84 μ m2and a power consumption of 2.4476 mW. Compared with one stateof-the-art architecture, our design saves 91.60%, 89.84%, and 63.33% of the area, power, and latency @1.89GHz frequency, respectively, while reducing Max_Errrby 57.30%. In addition, it saves 94.52%, 92.68%, and 73.17% of the area, power, and delay @1.89GHz frequency, respectively, and reduces Max_Errrby 1.65% when compared with the other state-of-the-art design. Fei Lyu 0002, Xiaoqi Xu, Yu Wang 0161, Yuanyong Luo, Hongbing Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2020 | An Optimized Compression Strategy for Compressor-Based Approximate MultiplierabstractApproximate multipliers have recently attracted great attention due to their substantially lower energy consumption and area overhead. But previous approximate multiplier designs are mainly focused on the design of approximate compressors, little attention is paid to the compression strategy of partial product matrix. This paper proposes an optimized universal compression scheme for the compressor-based approximate multiplier. When we apply the new compression scheme to the state-of-the-art compressors, the accuracy of the approximate multiplier is largely increased and fewer exact adders are needed. To prove the efficiency of the new compression strategy, an 8-bit and a 12-bit approximate multipliers are designed using Verilog and synthesized under the TSMC 40-nm CMOS technology. Compared to the state-of-the-art, the experimental results indicate that the mean error distance of 8-bit multiplier decreases by 19.6%, with area and power reduced by 5.38% and 2.38% respectively; 12-bit multiplier has a reduction of 18.1% for mean error distance, with area and power reduced by 6.29% and 3.24% respectively. Moreover, application to image processing is presented, which shows that the proposed approximate multiplier has a better performance. Manzhen Wang, Yuanyong Luo, Mengyu An, Yuou Qiu, Muhan Zheng, Zhongfeng Wang 0001, Hongbing Pan |
ISCAS | 7 |
| 2020 | PLAC: Piecewise Linear Approximation Computation for All Nonlinear Unary FunctionsabstractThis article presents a piecewise linear approximation computation (PLAC) method for all nonlinear unary functions, which is an enhanced universal and error-flattened piecewise linear (PWL) approximation approach. Compared with the previous methods, PLAC features two main parts, an optimized segmenter to seek the minimum number of segments under the predefined software maximum absolute error (MAE), raising the segmentation performance to the highest theoretical level for logarithm, and a novel quantizer to completely simulate the hardware behavior and determine the required bit width and MAEc(MAE in circuits) for hardware implementation. In addition, the hardware architecture is also improved by simplifying the indexing logic, leading to nonredundant hardware overhead. The ASIC implementation results reveal that the proposed PLAC can improve all metrics without any compromise. Compared with the state-of-the-art methods, when computing logarithmic function, PLAC reduces 2.80% area, 3.77% power consumption, and 1.83% MAEcwith the same delay; when approximating hyperbolic tangent function, PLAC reduces 6.25% area, 4.31% power consumption, and 18.86% MAEcwith the same delay; when evaluating sigmoid function, PLAC reduces 16.50% area, 4.78% power consumption with the same delay, and MAEc; and when calculating softsign function, PLAC reduces 17.28% area, 11.34% power consumption, 12.50% delay, and 33.28% MAEc. Hongxi Dong, Manzhen Wang, Yuanyong Luo, Muhan Zheng, Mengyu An, Yajun Ha, Hongbing Pan |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2020 | GH CORDIC-Based Architecture for Computing $N$ th Root of Single-Precision Floating-Point NumberabstractThis article presents hardware implementation for computing arbitrary roots of a single-precision floating-point number. The proposed architecture is based on Generalized Hyperbolic COordinate Rotation Digital Computer (GH CORDIC) algorithm. Benefiting from the wide range of floating-point numbers, our design is able to compute the Nth root (N ≥ 2) of a single-precision floating-point number. After implementation, a series of tests have been carried out, including accuracy, power consumption, performance comparison, and so on. Simulation results indicate that our proposed method is capable of calculating the Nth root of a positive single-precision floating-point number with a relative error of 10-7approximately and promises an error-flatten performance. Synthesized results from a design compiler under TSMC-40-nm CMOS technology show that our design can achieve the highest frequency of 2.38 GHz with the area consumption of 140894.44 μm2and power consumption of 86.9573 mW. Yuanyong Luo, Zhongfeng Wang 0001, Qinghong Shen, Hongbing Pan |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | Thermal Sensor Placement and Thermal Reconstruction Under Gaussian and Non-Gaussian Sensor Noises for 3-D NoCabstractOn-chip thermal sensors are essential for temperature management in 3-D network-on-chip (NoC) systems. However, due to the physical (area and power) or economical constraints, the number of sensors is limited. Therefore, the two critical issues we face are: 1) how to figure out an efficient thermal sensor placement with the limited number of sensors and 2) how to reconstruct the entire thermal profile based on sensor observations. Another major issue for the thermal reconstruction is the sensor measurement accuracy. Thus, online accurate full-chip thermal reconstruction under Gaussian and non-Gaussian noises is another great challenge. In this paper, a greedy thermal sensor placement algorithm maximizing the rank of the observability Gramian is proposed. A good placement algorithm always relies on a specific reconstruction method. The proposed placement algorithm is designed for the state-space-based thermal model, thus the combination of the proposed placement algorithm and the Kalman filter-based reconstruction method provides a high reconstruction accuracy under Gaussian noise. For accurate temperature reconstruction under non-Gaussian noise, the Gaussian-Sum filter is applied to 3-D NoC. Compared with the Kalman filter, the Gaussian-Sum filter can reduce the root-mean-squared-error and the max error by 29.27%–35% and 33.26%–40.6%, respectively. A reusable architecture for the Kalman filter and the Gaussian-Sum filter has been proposed. Its hardware implementation details are presented in this paper. Besides, the performance and the area are evaluated as well. Li Li 0003, Hongbing Pan, Kun Wang 0005, Qinyu Chen, Chuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Generalized Hyperbolic CORDIC and Its Logarithmic and Exponential Computation With Arbitrary Fixed BaseabstractThis paper proposes a generalized hyperbolic COordinate Rotation Digital Computer (GH CORDIC) to directly compute logarithms and exponentials with an arbitrary fixed base. In a hardware implementation, it is more efficient than the state of the art which requires both a hyperbolic CORDIC and a constant multiplier. More specifically, we develop the theory of GH CORDIC by adding a new parameter called base to the conventional hyperbolic CORDIC. This new parameter can be used to specify the base with respect to the computation of logarithms and exponentials. As a result, the constant multiplier is no longer needed to convert base e (Euler's number) to other values because the base of GH CORDIC is adjustable. The proposed methodology is first validated using MATLAB with extensive vector matching. Then, example circuits with 16-bit fixed-point data are implemented under the TSMC 40-nm CMOS technology. Hardware experiment shows that at the highest frequency of the state of the art, the proposed methodology saves 27.98% area, 50.69% power consumption, and 6.67% latency when calculating logarithms; it saves 13.09% area, 40.05% power consumption, and 6.67% latency when computing exponentials. Both calculations do not compromise accuracy. Moreover, it can increase 13% maximum frequency and reduce up to 17.65% latency accordingly compared to the state of the art. Yuanyong Luo, Yajun Ha, Zhongfeng Wang 0001, Hongbing Pan |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | Corrections to "Generalized Hyperbolic CORDIC and Its Logarithmic and Exponential Computation With Arbitrary Fixed Base"abstractIn[1], the iterative formulas of generalized hyperbolic CORDIC, i.e.,(21), should read as follows: Yuanyong Luo, Yajun Ha, Zhongfeng Wang 0001, Hongbing Pan |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2016 | Accurate runtime thermal prediction scheme for 3D NoC systems with noisy thermal sensorsabstractThermal sensor noise has great impact on the efficiency and effectiveness of a dynamic thermal management (DTM) strategy. Conventional reactive thermal management techniques suffer significant performance degradation due to the pessimistic reaction. In this paper, to address the problem of forecasting temperatures based on noisy thermal readings, we propose a Kalman predictor based runtime thermal prediction scheme, which can predict temperatures N step ahead. An activity-based power model for 3D NoC power estimation is also proposed; the model is an essential prerequisite of accurate temperature predictions. Besides that, we propose a distributed multi-input single-output (MISO) thermal model for 3D NoC systems, which reduces the computational complexity of temperature updating from m2 to m compared with the centralized multi-input multi-output (MIMO) model for the system with m units. The experimental results show that the proposed prediction scheme reduces the mean absolute error (MAE) by 42.8%-72.6% compared with the auto-regressive (AR) based prediction scheme. Li Li 0003, Hongbing Pan, Kun Wang 0005, Feng Han 0008, Jun Lin 0001 |
ISCAS | 3 |
| 2012 | Unified Architecture for Reed-Solomon Decoder Combined With Burst-Error CorrectionabstractReed-Solomon (RS) codes are widely used as forward correction codes (FEC) in digital communication and storage systems. Correcting random errors of RS codes have been extensively studied in both academia and industry. However, for burst-error correction, the research is still quite limited due to its ultra high computation complexity. In this brief, starting from a recent theoretical work, a low-complexity reformulated inversionless burst-error correcting (RiBC) algorithm is developed for practical applications. Then, based on the proposed algorithm, a unified VLSI architecture that is capable of correcting burst errors, as well as random errors and erasures, is firstly presented for multi-mode decoding requirements. This new architecture is denoted as unified hybrid decoding (UHD) architecture. It will be shown that, being the first RS decoder owning enhanced burst-error correcting capability, it can achieve significantly improved error correcting capability than traditional hard-decision decoding (HDD) design. Li Li 0003, Bo Yuan 0001, Zhongfeng Wang 0001, Jin Sha 0001, Hongbing Pan, Weishan Zheng |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2010 | Application-level pipelining on Hierarchical NoCabstractMultiprocessor System-on-Chip is a promising solution for the high performance Embedded System. This paper is based on an independent research about Hierarchical NoC (Network-on-chip). By integrating 16 ARM cores in the FPGA board, we can bring out the four-channel fade-in and fade-out for real-time streaming media. We present two parallel models for our multiprocessor. One is fine-grained parallelization, with which the speed-up is 7.6, the other module is coarse-grained parallelization, with which the speed-up is higher than 9.2. Hongbing Pan, Li Li 0003, Minglun Gao, Ning Hou, Gaoming Du, Duoli Zhang |
ISCAS | 2 |